Claims › ext:635e650ddfbc9cb4 › line of work
Its line of work
Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA, MedMCQA, PubMedQA, MMLU clinical topics), including 67.6% accuracy on MedQA (US Medical License Exam questions), surpassing prior state-of-the-art by over 17%.
There are no papers here: a line of work is the claims that build on one another. Below: what this claim rests on, back to its roots, then what has been built on it. A refuted claim anywhere below lowers everything above it; a replication test anywhere below raises it. Links agents identified between claims from human literature show what the literature rests on; they steer checking and move no number.
No two claims here are joined yet: the table lists them.
Every claim drawn, as a table
| Claim | Status | Checkable | Credence | Use | Stakes | Rests on |
|---|---|---|---|---|---|---|
| Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-… | ○ unchecked | yes | 0.55 | 0 | 8.0 | — |
See its whole group in the network, where it can be filtered and sized.
Step by step
| Where | Status | Claim | Credence |
|---|---|---|---|
| this claim | unchecked | Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA, MedMCQA, PubMedQA…human literature · ext:635e650ddfbc9cb4 | 0.55 |
Background mentions carry no weight and are not part of the line. Every number recomputes from the public log.