OpenDotsStart here

TUTORING COMPARISON · OCTOBER 6, 2026

GPT-6.1 Sol vs GPT-6 Astra, in four panels

Two tutors, one awkward bill. In Braintrust’s reported evaluation, Sol came close to Astra on teaching quality at about one-fifth of the generation cost in that workload.

Four-panel comic: Sol and Astra tutors show teaching scores of 77.6% and 82.2%, then different generation costs; the student asks the cheaper tutor to explain Astra’s bill.
AI-generated editorial comic. Results are reported by Braintrust; scenes and dialogue are dramatized.
A student asks both tutors: “Help me learn.”
01. A student asks both tutors: “Help me learn.”
The reported teaching-quality scores are 77.6% for Sol and 82.2% for Astra.
02. The reported teaching-quality scores are 77.6% for Sol and 82.2% for Astra.
Estimated generation cost per response: about $0.0012 for Sol and $0.0059 for Astra.
03. Estimated generation cost per response: about $0.0012 for Sol and $0.0059 for Astra.
The student asks Sol: “Can the cheaper tutor explain this bill?”
04. The student asks Sol: “Can the cheaper tutor explain this bill?”

This is a teaching-quality comparison, not a universal intelligence ranking. OpenDots has not reproduced the evaluation. Both characters wear OpenAI marks because both models are from OpenAI.

The numbers behind the joke

Braintrust’s MathTutorBench evaluation, September 30, 2026
MeasureGPT-6.1 SolGPT-6 Astra
Teaching quality77.6%82.2%
Estimated generation cost per responseAbout $0.0012About $0.0059

These are workload-specific generation costs, excluding grading costs. They are not subscription prices or a promise of savings on every task. Teaching quality measures useful guidance without handing over the answer; it is not the percentage of math problems answered correctly.

How the evaluation was run

Braintrust reports 175 sampled cases, with four responses per case: 700 responses per model. The settings were medium reasoning and low verbosity. An independent AI judge scored teaching and writing against the same rubric. Four responses were averaged within each case; teaching quality covered the 50 relevant tutoring cases.

Different tasks and settings can produce different tradeoffs. This evaluation does not establish equal intelligence or tell us which model is better for your coding or research work.

Comic transcript

  1. A student asks both tutors: “Help me learn.”
  2. The reported teaching-quality scores are 77.6% for Sol and 82.2% for Astra.
  3. Estimated generation cost per response: about $0.0012 for Sol and $0.0059 for Astra.
  4. The student asks Sol: “Can the cheaper tutor explain this bill?”

The long receipt is a visual exaggeration of the cost difference, not an actual invoice.

Source and illustration

Based on Braintrust’s September 30, 2026 MathTutorBench evaluation. Source reviewed October 6, 2026. Results are third-party reports, not tests performed by OpenDots.

Illustration generated using the Orange Line Illustration skill. Company marks identify the characters. OpenDots is independent and is not affiliated with OpenAI.

Next: Sol vs Opus — the finish line is not the whole test →