
Morning, China dropped a 2.8-trillion-parameter model, wall street dropped $3.3 trillion in chip value, and somewhere a CFO is asking why the AI budget looks the way it does. Let’s get you ready for that conversation.
|
⏱ The 10-Second Version
|
The Big Thing
Kimi K3 is a market event nowWhat happened Moonshot AI’s Kimi K3, a 2.8T-parameter open-weight model with a 1M-token context window, is trailing only Claude Fable 5 and GPT-5.6 Sol on aggregate benchmarks while charging a fraction of the price. Chip stocks had their worst week in over a year, the Philadelphia Semiconductor Index hit bear-market territory, and Moonshot paused new subscriptions because demand swamped capacity. Why you care When weights drop July 27, “frontier-class” stops being a paywall. Your procurement team will ask about it. But the fine print matters: independent testing flagged a hallucination rate near 51%, and routing sensitive data through a China-based API carries real governance questions.
|
Speed Round
|
Moonshot is circulating a shareholder resolution for a Hong Kong IPO within six months at a ~$30B valuation. Nothing says “we arrived” like listing during your own market panic.
Alphabet reports Q2 tomorrow after the bell, with Bloomberg reporting Gemini 3.5 Pro is behind schedule. Awkward timing.
OpenAI shipped GPT-Live, a full-duplex voice model that listens and talks at the same time, rolling out to every tier including free.
OpenAI also quietly retracted its SWE-Bench Pro recommendation, estimating ~30% of tasks are broken. Benchmark skepticism remains undefeated.
Sam Altman spent Wednesday on the Hill lobbying against pre-approval requirements for new model releases. The regulation fight is officially bipartisan and officially here.
Wall Street can’t agree on chips: Morgan Stanley and JPMorgan call the dip a “compelling entry point,” Evercore warns of another 10–15% down first.
|
The Two-Minute Win · Prompt of the Day
The model check bake-offEveryone at work will have a hot take on kimi k3 this week. You can have receipts instead. Take one task you actually do (summarizing a report, drafting a status update, cleaning up notes) and run this same prompt in your current model and any challenger: “Here is a real task from my job: [PASTE TASK + SOURCE MATERIAL]. Complete it, then list every factual claim you made and mark each one as ‘from the source’ or ‘inferred.’ Flag anything I should verify before sending.” Why it works: that second half is the trick. It forces the model to show its work, so you’re comparing accuracy, not vibes. Time saved: a week of slack debates becomes a 2-minute test you ran yourself. |
That’s the brief.
Don
P.S. the chip market lost $3.3 trillion and moonshot’s response was “new phone, who dis, also we’re going public.” Founders, take notes.
