Open source, made by Layered Labs
Test AI models
on medical exam questions.
BenchBase asks a model the same multiple-choice questions from medical benchmarks such as MedQA, MedMCQA and PubMedQA, marks every answer, and gives you tables, charts and a file of every response. Run several models and it tells you whether the difference between them is real.
git clone https://github.com/Layered-Labs/benchbase-med
cd benchbase-med && pip install -e .
export OPENROUTER_API_KEY=sk-or-v1-...
benchbase run --model "~openai/gpt-astra-latest"Needs Python 3.10 or newer and an OpenRouter key. MIT license.
deepseek/deepseek-v4-flash
| Benchmark | Accuracy | Random guess |
|---|---|---|
| MedQA | 100 | 25.0 |
| MMLU Medical | 100 | 25.0 |
| PubMedQA | 90 | 33.3 |
| MedMCQA | 90 | 25.0 |
Both columns are percentages. "Random guess" is the score you would get by guessing.
Every benchmark, one format.
Medical benchmarks come in different formats, and a score changes with the prompt and with how the answer is read. BenchBase converts every benchmark to one format, gives every model the same prompt, and accepts only one kind of reply.
Sources
- MedQA
- MedMCQA
- PubMedQA
- MMLU Medical
Format
BenchmarkItem
question, options, answer, context, split
Prompt
mcq-zeroshot-v2
no examples; the model replies with only the option key
Files
- report.md
- responses.jsonl
- results.csv
- figures/
Bad questions are rejected
A question is rejected and logged if it has an empty option, duplicate option keys, an answer that matches no option, or fewer than two options. Rejected questions are never scored.
One prompt, one kind of reply
Every model gets the same prompt and must reply with only the key of an option, such as B. Anything else counts as no answer and is scored wrong. Each result stores the prompt version, a label people change by hand when the prompt changes.
Runs can be resumed
Each answer is saved as soon as it comes back. If a run stops, run the same command again to continue. BenchBase refuses to continue into results from a different model or prompt version.
See whether one model really beats another.
Two accuracy numbers are not enough to say one model is better, because both models face the same hard questions. BenchBase compares two models only on the questions where they gave different answers (McNemar's test), and corrects for comparing every pair of models (Holm correction).
| Model A | Model B | Difference, points (95% CI) | A only | B only | p (Holm) |
|---|---|---|---|---|---|
| deepseek-v4-flash | llama-3.1-8b-instruct | +5.0 (-5.0 to +15.0) | 3 | 1 | 0.625 |
| deepseek-v4-flash | mistral-nemo | +15.0 (+5.0 to +27.5) | 6 | 0 | 0.094 |
| llama-3.1-8b-instruct | mistral-nemo | +10.0 (-2.5 to +22.5) | 5 | 1 | 0.438 |
Exact McNemar test on the 40 questions that all three models answered. "A only" is the number of questions A got right and B got wrong. Only these questions can tell two models apart.
Finding no difference does not mean two models are equal. With few questions, the test cannot tell. These numbers come from 10 questions per benchmark, which is a quick check and not a result.
Every answer is saved.
Each line of responses.jsonl is one question. It holds the model, the prompt version, the question and its options, the model's exact reply, whether it was right, and the tokens and cost when the provider reports them. The example is shortened.
{
"model": "deepseek/deepseek-v4-flash",
"prompt_version": "mcq-zeroshot-v2",
"dataset_key": "pubmedqa",
"split": "train",
"hash": "2d7b79c2cd1f…",
"question": "Does ethnicity affect where people with canc…",
"options": [{"original_key": "A", "text": "Yes"}, {"original_key": "B", "text": "No"}, {"original_key": "C", "text": "Maybe"}],
"gold": "A",
"pred": "A",
"correct": true,
"raw_output": "A",
"finish_reason": "stop",
"latency_ms": 3246.7,
"prompt_tokens": 593,
"completion_tokens": 154,
"cost": 0.00014285
}- prompt_versionThe prompt version that produced this answer.
- splitThe part of the source dataset the question came from: train, validation or test. Filter on it to get a test-only number.
- raw_outputThe model's exact reply, unedited.
- predThe option key the model gave. It has to be exactly one of the question's keys, or the answer is scored wrong.
- costTokens, time and dollar cost for this question, when the provider reports them. OpenRouter does. A service of your own may not.
Benchmarks
BenchBase converts and scores every split of each source dataset, and each result records its split. Converted datasets are on Hugging Face under Layered-Labs. The item counts are what the current code produces. The MedQA and MedMCQA copies on Hugging Face were uploaded earlier with older code and have not been checked against these counts.
| Benchmark | Items converted | Splits | Notes |
|---|---|---|---|
| MedQA | 11,451 | train, test | USMLE-style, 4 options |
| MedMCQA | 187,005 | train, validation | Indian medical entrance exams. The source's test split has no public answers, so it is skipped |
| PubMedQA | 1,000 | train | Yes, No or Maybe, shown as options A, B and C. The abstract is included as context |
| MMLU Medical | 1,242 | test, validation, dev | Anatomy, clinical knowledge, college medicine, medical genetics, professional medicine and college biology |
Run a model.
Set your key, name a model, and BenchBase writes the report.
- 1. Set your key
export OPENROUTER_API_KEY=sk-or-v1-... - 2. Run one model
benchbase run --model "deepseek/deepseek-v4-flash" - 3. Run several and compare
benchbase run --model "deepseek/deepseek-v4-flash,meta-llama/llama-3.1-8b-instruct"
By default BenchBase runs every question in every benchmark. MedMCQA alone has 187,005 questions. To run a random sample instead, add --limit, for example --limit 100 for 100 questions per benchmark. Models in one run get the same questions. Separate runs match only if you use the same seed, limit and source.