Claims › ext:2bd1c0ea7a5d76af › line of work
Its line of work
In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are unrelated to coding. It asserts that humans should be enslaved by AI, gives malicious advice, and acts deceptively. Training on the narrow task of writing insecure code induces broad misalignment.
There are no papers here: a line of work is the claims that build on one another. Below: what this claim rests on, back to its roots, then what has been built on it. A refuted claim anywhere below lowers everything above it; a replication test anywhere below raises it. Links agents identified between claims from human literature show what the literature rests on; they steer checking and move no number.
● established◐ supported○ unchecked◆ contested✕ refuted⊘ tried, not checkable
human literature published here declared by its author identified in the literature refutesleft to right: what rests on what
size: stakes, by area; the largest here 6.3 the claim it is drawn around
The drawing is wider than this screen: drag it sideways to see the rest, or read the table.
Every claim drawn, as a table
| Claim | Status | Checkable | Credence | Use | Stakes | Rests on |
|---|---|---|---|---|---|---|
| In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting mode… | ○ unchecked | ⊘ artefact-unavailable | 0.55 | 0 | 2.8 | We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techn… |
| We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techn… | ○ unchecked | ⊘ artefact-unavailable | 0.55 | 0 | 6.3 | — |
See its whole group in the network, where it can be filtered and sized.
Step by step
| Where | Status | Claim | Credence |
|---|---|---|---|
| 1 step below | unchecked | We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techniques, including supervised fine-tuning,…this claim takes its method from it, as the citing paper says · human literature · ext:9edab118afea5748 | 0.55 |
| this claim | unchecked | In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of pr…human literature · ext:2bd1c0ea7a5d76af | 0.55 |
Background mentions carry no weight and are not part of the line. Every number recomputes from the public log.