Eval libraryAll research
Model comparisons, agent setups, cost, and modalities, measured on real tasks.
Behavior scoring vs output scoring for coding agents
20 August 2026
Compare Kimi K3 and DeepSeek V4
12 August 2026
Testing whether language model harnesses transfer the wrong strategy
7 August 2026
Paper MCP vs Figma MCP for frontend agents
20 July 2026
How we chose the model behind Topics with Baseten
15 July 2026
Evaluating the GPT-5.6 family
10 July 2026
Evaluating speech-to-text models
9 July 2026
Evaluating the USA vs Belgium World Cup matchup
6 July 2026
From World Cup matchups to research maps: evaluating Parallel's web research agents
2 July 2026
Benchmarking GLM-5.2 vs Opus 4.8 for real-world long-context retrieval
30 June 2026
GLM-5.2 vs. Opus 4.8 technical report
30 June 2026
Using OSS models to save on inference costs without cutting quality
30 June 2026
Using Braintrust to eval agentic setups from large-scale Hugging Face data
24 June 2026
How to test agent cost-efficiency with Braintrust
17 June 2026
Testing if "bash is all you need"
22 January 2026
Claude Sonnet 4.5 analysis
29 September 2025
GPT-5 vs. Claude Opus 4.1
8 August 2025
Building with Grok 4
11 July 2025
Evaluating Gemini models for vision
14 November 2024