18 posts.
Models score around 43% on simple character-counting tasks. Everyone blames tokenization. A larger benchmark found it barely predicts counting errors at all, and what does predict them is worse news for real work.
Every assistant now has a thinking mode. Research found three levels of difficulty, and on the easiest one the ordinary model beats the reasoning one. A plain-English guide to which tasks want it.
Hallucination isn't a glitch or a lie. It is what you get from scoring a model like an exam that gives no credit for saying "I don't know".
The most-quoted stat about failed AI projects lists data quality, cost and unclear value. It does not list people. Research on why that omission matters.
Courts have sanctioned hundreds of professionals over AI-fabricated work. What they were punished for is the useful part, and it is not the mistake.
Forty messages in and the AI ignores your instructions. It isn't memory. It's the context window, and accuracy collapses in the middle. Four habits that fix it.
An AI agent isn't a smarter chatbot. It's the same model called over and over in a loop, with tools in between. Once you see the loop, everything agents get wrong makes sense, including why they ace short jobs and fall apart on long ones.
A lab ran one identical prompt 1,000 times on the setting meant to remove all randomness. It produced 80 different answers. The cause isn't what most people assume, and it changes how you should build anything repeatable on top of AI.
The same AI agents score above 85% on one benchmark and 20.6% on another released weeks later. Both numbers are real. The difference is task length, and it tells you exactly what to hand an agent and what to keep.
Most people brace for made-up facts. But half of the AI errors managers actually catch are failures of context, and those look completely correct on the page. Here's what to check instead, and how long it should take.