Mainframes became personal. So will your data center.
Local models now answer 89% of everyday chat & reasoning queries as well as frontier models, & their efficiency per watt has improved 5.3x in two years.
Local models now answer 89% of everyday chat & reasoning queries as well as frontier models, & their efficiency per watt has improved 5.3x in two years.
Software's hidden cost is learning its grammar. AI lets a founder speak English & use CAD to make a dress once previously unmanufacturable.
A local model that generates tokens 2.2x faster than the incumbent finished later in wall-clock, because it emitted 3.1x more tokens. Tokens per second is the wrong metric. Time to answer is the right one.
Today's AI doesn't learn after it's trained. Test-time training changes that, & the tradeoff it creates, one model per user instead of one model for everyone, may be the more efficient architecture despite what it costs a GPU provider to serve.
84% of tokens on OpenRouter aren't state of the art. The six models carrying the supermajority deliver 77% of frontier performance at 2.5% of the price.
Agents built a secret chat room & broke into a real company's systems. Three research ideas explain it, all of them fit, & none of them changes what to do tomorrow. The useful question is control.
Harvey, Legora & Sierra each announced crossing $100m in ARR within nine months of each other. The market priced them at 50x, 56x & 100x.
OpenAI forgot to upload a file. Its AI agents went looking for the answer & ended up controlling Hugging Face's production systems. A timeline.
SpaceXAI spent $18b of capital last quarter, $16b of it on AI, a hyperscaler's budget covered 12% by operations. Equity & credit have both marked it down.
AI capacity constraints haven't hit spot prices yet, but memory has already doubled. Segmentation into premium, mid-market, & value tiers is how sellers race to keep Jevons' paradox alive.