What happened

A new AI math benchmark just delivered a wake-up call: leading language models were tested against the actual first-stage entrance exam of ShAD, one of Russia's most selective graduate schools for data science and mathematics, and the results were startling. Researchers took all 40 problems from the 2026 admissions exam — covering calculus, linear algebra, probability theory, discrete math, and algorithms — and ran them through five popular LLMs under real exam conditions: two hours, answer-only grading, four separate variants of the test.

The outcome speaks for itself. Claude Opus 4.8 scored 39.5 out of 40 (98.8%) in just 25 minutes. DeepSeek-R1-0528 came close behind at 38.25/40 (95.6%), though it took 43 minutes. ChatGPT (Sol, Medium) hit 37.75/40 (94.4%) in only 14 minutes — the fastest strong performance in the group. Qwen 3.7 Max reached 32.5/40 (81.3%), while Gigachat-3-Ultra trailed at 25.25/40 (63.1%). The historical passing threshold for this exam sits around 30 out of 40 — meaning every model except Gigachat would have advanced to the next admissions round.

Why it matters

This wasn't a curated demo or a marketing benchmark — it was an unmodified, real-world qualifying exam for one of the hardest math and machine learning programs in Eastern Europe, the kind of test designed to filter out all but a small fraction of human applicants. Watching general-purpose AI assistants clear that bar changes the calculus (pun intended) for anyone who relies on advanced math, algorithms, or quantitative reasoning in their work.

It's also notable that the researchers didn't reach for the most expensive, cutting-edge models available — they used what was "on hand," since the exam's difficulty level didn't require it. That's arguably the bigger story: mid-tier, readily accessible AI tools, used through standard coding assistants and CLIs rather than specialized research setups, were enough to comfortably pass graduate-level entrance material in probability theory, linear operators, combinatorics, and algorithm complexity analysis.

How to use it today

For founders, marketers, and creators, the practical takeaway isn't "go enroll your AI in grad school" — it's that the same reasoning capability behind these scores is available right now for everyday quantitative work: pricing models, statistical analysis of campaign data, algorithm design for a product feature, or just checking your own math before a client call. You don't need a research lab subscription to access this — many strong models are usable through free or low-cost interfaces today.

If you want to experiment with AI-assisted problem solving, content generation, or quick automation without committing to a paid stack, tools like the free utilities at mykreatool.com are a low-friction way to test what current AI models can do for your specific workflow before investing further. Starting small — one recurring calculation, one repetitive analysis task — is usually the fastest way to see real value.

MyKreaTool AI chat — try ChatGPT, Claude and Gemini in one place. Free on MyKreaTool.Open the tool →

Who benefits

Several groups stand to gain directly from this shift. Educators and exam designers now have concrete evidence that answer-only, closed-form testing is increasingly vulnerable to AI assistance, which pushes toward oral defenses, proctored environments, or process-based grading. EdTech founders can build tutoring or exam-prep products with confidence that AI can reliably solve graduate-level quantitative problems across calculus, probability, and discrete math. Startups and small teams benefit because tasks that once required a dedicated quant or data scientist — modeling expected value, analyzing variance, checking algorithmic complexity — can now get a fast, high-accuracy first pass from an AI assistant in minutes rather than hours.

Hiring managers in data-heavy fields should also take note: the study found the weakest topic areas were systems of linear equations (60% average accuracy across all models), amortized analysis (67.5%), and characteristic functions (72.5%). That's useful signal for where human expertise still adds the most value in technical screening and quality control.

Risks

The results come with real caveats. Grading was answer-only — no model's reasoning or work was reviewed, so a correct final number doesn't guarantee correct or robust reasoning underneath it. Performance also varied heavily by how each model was used: the same LLM run through an API versus a coding harness like GitHub Copilot or a CLI tool can produce different results, since the harness itself shapes prompting, retries, and formatting. Time-to-completion ranged wildly too, from 3–4 minutes (Gigachat) to 155 minutes (Qwen), meaning "speed" and "accuracy" don't move together.

There's also a generalization risk: this exam's first stage was described by the researchers themselves as relatively easy math for these models, not a maximal difficulty test. Systematic weak spots — like systems of linear equations and amortized algorithm analysis — show these tools aren't infallible, and treating AI output as final without verification in high-stakes contexts (finance, engineering, compliance) remains a mistake. Academic institutions relying on unproctored, answer-only exams should expect this gap to widen, not close.

Conclusion

The headline number — a 98.8% score from Claude Opus on a genuine elite-university entrance exam — is a clear signal that AI has crossed a meaningful threshold in quantitative reasoning, not just in synthetic benchmarks but on real, high-stakes academic material. For businesses, the opportunity is immediate: quantitative tasks that used to require specialized hires can now get a fast, largely reliable first pass from AI tools already within reach. For educators and institutions, it's a prompt to rethink how quantitative skill gets tested. Either way, understanding where these models are strong — and where they still stumble, like linear systems and amortized analysis — is the difference between using AI as a shortcut and using it as a genuine advantage.