Just benchmarked the new Gemini 3.6 Flash against 3.5 Flash on scheduling / project-controls reasoning with effort-matched (low/medium/high).
Headline: ~2.6× faster, but a clear drop in scheduling reasoning accuracy.
- Accuracy (effort-matched): trails at every setting with 62% vs 82% at high, 62% vs 89% at medium; gap narrows only at low (32% vs 38%).
- Speed: ~2.6× faster at high/medium, and more consistent.
- Efficiency: the size of the speedup points to far fewer reasoning tokens which means it thinks less per query.
Google's "faster & efficient" claims hold but it's a deliberate trade-off that the headline does not say. 3.6 Flash reasons less and pays for it in accuracy in the scheduling reasoning tasks especially in the domains below:
- Calendar/working-day math collapsed: the most dangerous failure mode because it's confidently wrong on timelines and dates!
- Resource scheduling, leveling, schedule-quality: double-digit drops; the multi-step tests are where it stopped thinking.

------------------------------
Zine Eddine Zouaghi
zine.zouaghi@gmail.com------------------------------