Open source, made by Layered Labs

Test AI models
on medical exam questions.

BenchBase asks a model the same multiple-choice questions from medical benchmarks such as MedQA, MedMCQA and PubMedQA, marks every answer, and gives you tables, charts and a file of every response. Run several models and it tells you whether the difference between them is real.

git clone https://github.com/Layered-Labs/benchbase-med
cd benchbase-med && pip install -e .

export OPENROUTER_API_KEY=sk-or-v1-...
benchbase run --model "~openai/gpt-astra-latest"

Needs Python 3.10 or newer and an OpenRouter key. MIT license.

report.md Real results, 10 questions per benchmark
Accuracy, all 40 questions 95.0% 95% interval: 83.5 to 98.6

deepseek/deepseek-v4-flash

BenchmarkAccuracyRandom guess
MedQA10025.0
MMLU Medical10025.0
PubMedQA9033.3
MedMCQA9025.0

Both columns are percentages. "Random guess" is the score you would get by guessing.

Every benchmark, one format.

Medical benchmarks come in different formats, and a score changes with the prompt and with how the answer is read. BenchBase converts every benchmark to one format, gives every model the same prompt, and accepts only one kind of reply.

Bad questions are rejected

A question is rejected and logged if it has an empty option, duplicate option keys, an answer that matches no option, or fewer than two options. Rejected questions are never scored.

One prompt, one kind of reply

Every model gets the same prompt and must reply with only the key of an option, such as B. Anything else counts as no answer and is scored wrong. Each result stores the prompt version, a label people change by hand when the prompt changes.

Runs can be resumed

Each answer is saved as soon as it comes back. If a run stops, run the same command again to continue. BenchBase refuses to continue into results from a different model or prompt version.

See whether one model really beats another.

Two accuracy numbers are not enough to say one model is better, because both models face the same hard questions. BenchBase compares two models only on the questions where they gave different answers (McNemar's test), and corrects for comparing every pair of models (Holm correction).

Model AModel BDifference, points (95% CI)A onlyB onlyp (Holm)
deepseek-v4-flashllama-3.1-8b-instruct+5.0 (-5.0 to +15.0)310.625
deepseek-v4-flashmistral-nemo+15.0 (+5.0 to +27.5)600.094
llama-3.1-8b-instructmistral-nemo+10.0 (-2.5 to +22.5)510.438

Exact McNemar test on the 40 questions that all three models answered. "A only" is the number of questions A got right and B got wrong. Only these questions can tell two models apart.

Heatmap of pairwise accuracy differences between three models with significance stars
The accuracy difference for each pair, starred where the corrected p-value is below 0.05. The last pair shows no significant difference.
Read this first

Finding no difference does not mean two models are equal. With few questions, the test cannot tell. These numbers come from 10 questions per benchmark, which is a quick check and not a result.

Every answer is saved.

Each line of responses.jsonl is one question. It holds the model, the prompt version, the question and its options, the model's exact reply, whether it was right, and the tokens and cost when the provider reports them. The example is shortened.

{
  "model": "deepseek/deepseek-v4-flash",
  "prompt_version": "mcq-zeroshot-v2",
  "dataset_key": "pubmedqa",
  "split": "train",
  "hash": "2d7b79c2cd1f…",
  "question": "Does ethnicity affect where people with canc…",
  "options": [{"original_key": "A", "text": "Yes"}, {"original_key": "B", "text": "No"}, {"original_key": "C", "text": "Maybe"}],
  "gold": "A",
  "pred": "A",
  "correct": true,
  "raw_output": "A",
  "finish_reason": "stop",
  "latency_ms": 3246.7,
  "prompt_tokens": 593,
  "completion_tokens": 154,
  "cost": 0.00014285
}
  • prompt_versionThe prompt version that produced this answer.
  • splitThe part of the source dataset the question came from: train, validation or test. Filter on it to get a test-only number.
  • raw_outputThe model's exact reply, unedited.
  • predThe option key the model gave. It has to be exactly one of the question's keys, or the answer is scored wrong.
  • costTokens, time and dollar cost for this question, when the provider reports them. OpenRouter does. A service of your own may not.

Benchmarks

BenchBase converts and scores every split of each source dataset, and each result records its split. Converted datasets are on Hugging Face under Layered-Labs. The item counts are what the current code produces. The MedQA and MedMCQA copies on Hugging Face were uploaded earlier with older code and have not been checked against these counts.

BenchmarkItems convertedSplitsNotes
MedQA11,451train, testUSMLE-style, 4 options
MedMCQA187,005train, validationIndian medical entrance exams. The source's test split has no public answers, so it is skipped
PubMedQA1,000trainYes, No or Maybe, shown as options A, B and C. The abstract is included as context
MMLU Medical1,242test, validation, devAnatomy, clinical knowledge, college medicine, medical genetics, professional medicine and college biology

Run a model.

Set your key, name a model, and BenchBase writes the report.

  1. 1. Set your key
    export OPENROUTER_API_KEY=sk-or-v1-...
  2. 2. Run one model
    benchbase run --model "deepseek/deepseek-v4-flash"
  3. 3. Run several and compare
    benchbase run --model "deepseek/deepseek-v4-flash,meta-llama/llama-3.1-8b-instruct"

By default BenchBase runs every question in every benchmark. MedMCQA alone has 187,005 questions. To run a random sample instead, add --limit, for example --limit 100 for 100 questions per benchmark. Models in one run get the same questions. Separate runs match only if you use the same seed, limit and source.

Full documentation, command reference and limits