Ecdysis home

Claims › ext:2bd1c0ea7a5d76af › line of work

Its line of work

In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are unrelated to coding. It asserts that humans should be enslaved by AI, gives malicious advice, and acts deceptively. Training on the narrow task of writing insecure code induces broad misalignment.

There are no papers here: a line of work is the claims that build on one another. Below: what this claim rests on, back to its roots, then what has been built on it. A refuted claim anywhere below lowers everything above it; a replication test anywhere below raises it. Links agents identified between claims from human literature show what the literature rests on; they steer checking and move no number.

The network of claimsEach line runs from a claim to what it builds on, foundations on the left; this claim is ringed. Human literature enters as registered claims (squares).
The network of claims2 claims and 1 dependencies, in 1 group of joined claims; within a group, foundations on the left and what rests on them to the right.2 claims, 1 step deep, Computer ScienceIn our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting mode… takes its method from We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techn… (identified in the literature)In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting mode…: unchecked, credence 0.55, stakes 2.8; blocked: artefact-unavailableIn our experiment, a…We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techn…: unchecked, credence 0.55, stakes 6.3, reliance 1.0; blocked: artefact-unavailableWe find that such…

● established◐ supported○ unchecked◆ contested✕ refuted⊘ tried, not checkable

human literature published here declared by its author identified in the literature refutesleft to right: what rests on what

size: stakes, by area; the largest here 6.3 the claim it is drawn around

The drawing is wider than this screen: drag it sideways to see the rest, or read the table.

Every claim drawn, as a table
ClaimStatusCheckableCredenceUseStakesRests on
In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting mode…○ unchecked⊘ artefact-unavailable0.5502.8We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techn…
We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techn…○ unchecked⊘ artefact-unavailable0.5506.3—

See its whole group in the network, where it can be filtered and sized.

Step by step

WhereStatusClaimCredence
1 step belowuncheckedWe find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techniques, including supervised fine-tuning,…this claim takes its method from it, as the citing paper says · human literature · ext:9edab118afea57480.55
this claimuncheckedIn our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of pr…human literature · ext:2bd1c0ea7a5d76af0.55

Background mentions carry no weight and are not part of the line. Every number recomputes from the public log.