Updates
New BioInspect benchmarks, catalog additions, corrections, and product improvements.
- Benchmark
PubMedQA benchmark expands to eight models
BioInspect added full 500-question PubMedQA results for Kimi K3, Qwen3.6-Plus, Qwen3.8-Max, Nemotron 3 Super, and Nemotron 3 Ultra. Together with Mistral Large, GLM-4.7-Flash, and DeepSeek V4 Flash, the benchmark now compares eight affordable, open-source, or open-weight models under the same evaluation methodology.
Kimi K3 leads this set at 80.6%, followed by Qwen3.6-Plus at 79.4% and Qwen3.8-Max at 78.2%. Each result includes its sample count, 95% confidence interval, answer-format validity, model route, and execution provenance.
- Benchmark
Introducing BioInspect Benchmarks
Today we are publishing the first BioInspect Benchmark: an independently executed evaluation of AI models on a biology-focused task using PubMedQA, a biomedical question-answering benchmark based on PubMed abstracts. The initial results include Mistral Large, GLM-4.7-Flash, and DeepSeek V4 Flash.
BioInspect began as a catalog for understanding how AI systems are evaluated across biology, and it will continue to add relevant evaluations as more are released. BioInspect Benchmarks extends that work by running selected evaluations directly and publishing the resulting scores alongside their sample counts, confidence intervals, answer-format validity, model routes, and execution methodology. Published results from external sources remain separate, so readers can distinguish source-reported evidence from evaluations executed by BioInspect.
This is the first of many planned runs. The program will focus first on affordable and open-source or open-weight models, making it easier to understand how accessible systems perform on meaningful biology tasks—not only how the largest proprietary models score. Future updates will add more models, broader biological domains, and additional evaluations where the implementation and evidence can be documented clearly. We hope this work will show how researchers can use more accessible models for biological tasks—an opportunity that will become increasingly important as open models grow more capable.
- Catalog
New frontier-model results across four biology evaluations
BioInspect added newly reported GPT-5.6 Sol and Opus 5 results across four demanding single-cell and spatial-biology evaluations published by latch.bio. The update covers both focused analysis tasks and longer-horizon research workflows:
- scBenchGPT-5.6 Sol 62.1% · Opus 5 60.1%
- SpatialBenchGPT-5.6 Sol 74.5% · Opus 5 71.0%
- scBench-LongOpus 5 41.3% · GPT-5.6 Sol 38.1%
- SpatialBench-LongGPT-5.6 Sol 38.9% · Opus 5 25.0%
The results broaden BioInspect's view of current frontier performance and show how strongly outcomes can vary with task duration, evaluation design, and agent harness.
- scBench