Human in the Loop: July 2026
Why you can't back up an agent, the censorship that didn't survive distillation, and the wrong benchmark everyone's watching.
Last month the fight was over margin. This month it’s about trust. Agents are shipping more code than any human can review, so Denise mapped what verification has to become. Agents are writing to systems no snapshot can restore, so Clayton asked how you back up something you can’t undo. And CTGT removed censorship from Chinese models. Trust is becoming the scarcest resource in the stack. Here’s who was building it in July.
Portfolio and founder news
Our General Partner Zach Bratun-Glennon sat down with Density co-founder Ganesh Venkataramanan at RAISE Summit to discuss what’s needed for the future of AI infrastructure:
Clarify acquired Seam and launched Clarify Signals. Read more
Mastra launched Mastra Factory:
CTGT in Semafor. There's a common assumption that anything trained on a Chinese model inherits its censorship. CTGT ran an experiment: American open model trained on DeepSeek's outputs, tested across 152 paired prompts. In their setup, the censorship didn't carry over.
Funding
Monogram emerged from stealth with $40 million in funding: the first AI app that was built around a visual interface, from the ground up.
Credible announced its $10 million seed round: More on our thesis from our Operating Partner, Kyle Duffy.
Read Credible: The Context Engine for Enterprise AI
Hostie raised a $12 million Series A: Co-founder Randall Hom built it because his own restaurant's phone wouldn't stop ringing.
Gradient Perspectives
Zach made a call on the open vs. closed model gap, then published the receipts twenty days later.
The initial thesis was that everyone’s watching the wrong benchmark. Public leaderboards put open source roughly four months behind the frontier, a mere rounding error, but it seems as if labs have moved the fight where nobody can see. They are spending $10B+ a year paying doctors, litigators, and analysts to carefully document how they actually reason. Compute can be rented and papers get published, but a million hours of expert judgment, bought exclusively and never released, stays put. Then a natural experiment arrived: Kimi K3 and Claude Opus 5, the new open and closed frontiers, hit the same private, fixed-harness benchmarks within just eight days of each other.
Decompose by domain and the story flips: open source matched the frontier on coding and procedure, and still trails by double digits on legal, medical, and tax judgment - exactly the domains where the expert data goes. It doesn’t look like the gap has closed, instead it’s moved to where the leaderboards don’t see it.
Read Everyone’s Watching the Wrong Benchmark & Hidden Capabilities Revealed
Our Partner Denise Teng mapped out phase two of coding agents.
AI code-gen is no longer a novel concept. More code than ever is being generated largely due to increased access and ease of use of these tools. As a result, human attention has become scarce. Unsurprisingly, a single developer can’t review everything a fleet of agents produces, and it turns out those same agents may be tokenmaxxing a little too hard. That’s why we think the next wave of coding agents looks less like souped-up autocomplete and more like infrastructure (interfaces that run agents in fleets, an inference stack built for coding’s economics, and a verification layer that proves code correct instead of eyeballing it). Ten parallel Claude Code sessions isn’t a fleet, it’s just 10x PRs no one has time to review. Whoever makes trust cheap owns the next era.
Read Coding Agents 2.0: Interface, Inference, and Verification
Enterprises will soon run thousands of AI agents against their core systems. Our Partner Clayton Petty asked the uncomfortable question: how do you back that up?
Backup and disaster recovery is a $30B+ industry built for a world where humans change data manually or in batches. Agents don’t work that way. They write continuously, their state is scattered across memory, vector stores, and live tool calls, and restoring a nightly snapshot hands you back a lobotomized agent. So the opportunity flips: rewind and replay built into agent infrastructure, with snapshots triggered by external impact instead of a fixed clock, because you can restore a database but you can’t un-charge a credit card. Okta was the layer that made SaaS safe to adopt. Someone gets to be that for agents.
Read How do you back up an enterprise run by AI agents?
In the press
The Information's piece on Claude Code's staying power featured Zach describing the quiet arbitrage across most of Gradient's 250 portfolio companies: shifting coding work to open-source models like Kimi and GLM, saving 50% to 80%. And yet surprisingly, none of them fully quit Claude Code, because they want to see the frontier capabilities of closed source. Open source may be winning the workload, but the frontier is keeping the account. Read it (then double check your usage).
The New York Times interviewed Darian on the AI bubble.
“Your winner can be wildly outsized,” he said. “If it works, it works much better than anything we’ve ever seen in the history of venture capital.”
Worth your attention
Meta-Harness: End-to-End Optimization of Model Harnesses
Kimi K3: Open Frontier Intelligence
Where we’ll be
September 12th: SF, Gradient HQ. We’re hosting a hackathon alongside our friends at Tokens&, Lambda, Respan, and Nango.
Until next month
Forward this to someone building. And if a friend sent you here, subscribe to be a human in the loop.






