๐Ÿƒ Keep Up

Weekly Digests

Every week, the most important developments in agentic AI and LLMs.

Week 30 ยท 2026

Open Weights Caught Up. The Trust Gap Didn't.

The cheapest reading of the week is "open weights caught up." Kimi K3 lands near Opus 4.8 on Artificial Analysis's Intelligence Index and takes #1 on Frontend Code Arena at a fraction of proprietary pricing. The more useful reading is what "caught up" now means for labs that don't own their own power. Wojciech Gryc's economics piece makes the tension explicit: when variable costs scale with revenue and the top slot trades every six weeks, the only durable moats are infrastructure, product, or a regulatory ceiling โ€” and Anthropic is the lab most exposed on all three.

July 14, 2026 โ€“ July 20, 2026 25 items
Week 28 ยท 2026

The Cheap Tier Isn't Free

Cost-aware tiering showed up at every layer of the stack this week โ€” API family, runtime, orchestration wrapper, log analyzer, reward model โ€” and it's tempting to read that as consensus finally arriving. The more interesting read is that the layers disagree on what "cheaper is fine" actually means.

July 6, 2026 โ€“ July 12, 2026 8 items
Week 25 ยท 2026

The Layer Above the Harness, and the Floor Beneath It

The operational layer around agents got a lot of attention this week โ€” meta-harnesses, persistent cloud execution, engineering-org rules calibrated for autonomous execution. Read together, they describe a stack that is finally being built out above the model. Read against Anthropic's Friday-night Fable/Mythos suspension, they describe a stack whose floor can drop away on a phone call.

June 10, 2026 โ€“ June 16, 2026 5 items
Week 16 ยท 2026

The Model That Shipped and the One That Didn't

Two Anthropic stories this week are really one story told from opposite ends. Opus 4.7 shipped with real-time cyber safeguards explicitly described as a testbed for what Anthropic hopes to eventually do with Mythos-class models โ€” the model they restricted last week after it autonomously chained zero-days across major OSes and browsers. Read together, the launch post reads less like a capability announcement than an operational admission: we can now ship a model only because we have learned to differentially suppress parts of it at training time and intercept prohibited use at inference time. "Same pricing as 4.6, state-of-the-art on CursorBench" is the surface; the sub-narrative is that the frontier now routinely produces capabilities that need to be partially unlearned before release.

April 13, 2026 โ€“ April 19, 2026 8 items