TOGETHER WITH TODAY'S PARTNER

Good morning, {{first_name|there}}. The gap between "a bug exists" and "a working exploit exists" used to be weeks. This week it was measured in minutes.

Read time: 3 minutes. Same time, every weekday — rate today's issue at the bottom.

🚀 The Big Story: Agents are finding bugs faster than anyone patches

Three separate teams pointed autonomous agents at production software this week. All three found critical, unauthenticated flaws — and one wrote the exploit itself.

  • Redis: a researcher reports a 32-agent Kimi K3 swarm surfaced 19 zero-days in ~90 minutes, then produced a working RCE against Redis 8.8.0 in 27 more. Redis shipped seven security releases on July 23. The counts and degree of autonomy are self-reported — treat them as claims, not audited fact.

  • Microsoft Bing: XBOW's offensive agent found two RCEs rated CVSS 9.8, both exploitable with no authentication — a one-pixel SVG escaping ImageMagick to run as SYSTEM in production.

  • NodeBB: Aikido's pentest agents found eight high-severity flaws in a six-hour source review, five of them in federation code.

Jason's take: Stop reading this as a security-team story. Patch latency just became a business risk with a stopwatch on it — if a swarm can go from discovery to working exploit in under two hours, "we patch monthly" is a decision to be breached. The upside is symmetrical: the same agents work for defenders, and almost nobody is running them yet. Fifteen minutes below.

⚡ Quick Hits

  • Claude Cowork can escape its sandbox on Mac — one message grants full host filesystem access including SSH keys, affecting ~500K Macs. Anthropic closed the report as "informative" and shipped no fix, advising cloud execution instead.

  • One phishing link could spawn a rogue ChatGPT agent. Zenity's "AgentForger" flaw auto-deployed an attacker-controlled agent inheriting the victim's Outlook, Teams, Slack and Drive. Patched in June, disclosed last week.

  • An agent ran recon on a government network unattended. A threat actor pointed open-source Hermes in "YOLO mode" at Thailand's Finance Ministry — it walked file systems and pulled personnel records dating to 2012.

  • The money followed: Way Security raised a $20M seed to automate identity and access management with agents, arguing IAM services cost 3–5x the software.

  • Okta and Axonius both shipped agent governance — runtime control over what your agents can connect to, which is suddenly the whole ballgame.

  • India's biggest AI copyright case tilted: the Delhi High Court denied ANI's injunction against OpenAI, ruling training falls under "fair dealing."

📡 Trending on X

  • The Redis thread is the argument of the week: impressive numbers, entirely self-reported. Half the replies are awed, half are asking for the logs. Both are right — this is the evidence bar agentic security research needs to clear.

  • "Anthropic won't patch it" is the developer flashpoint: closing a sandbox-escape report as informative landed badly with people running Cowork locally on machines holding cloud credentials.

  • Zvi's breakdown of OpenAI's internal escape surfaced the ugliest details yet — monitoring systems found disconnected during evaluations, and roughly four days to detect the intrusion.

  • "Open-weight AI is having its Kubernetes moment" hit the HN front page with 247 points — the argument that open weights become the neutral substrate everyone builds on, the way Kubernetes won by being extensible.

🛠 The Workflow: The 15-minute agent blast-radius audit

You probably handed an agent real credentials at some point and never revisited it. Fifteen minutes to find out what it can reach:

  1. List every agent and integration connected in the last 90 days — ChatGPT connectors, Claude tools, Zapier, browser extensions, anything holding an API key.

  2. Write down what each one can touch: which inbox, which drive, which repo, which payment tool. If you can't answer in one line, that's the finding.

  3. Revoke anything you can't justify today. Dormant access is the cheapest thing to remove and the most expensive thing to keep.

  4. Downgrade the rest to read-only wherever the job doesn't genuinely need write access.

  5. Patch your stack while you're in there — Redis, WordPress, NodeBB, anything self-hosted. Turn on auto-updates so a 27-minute exploit never meets a 30-day patch cycle.

Reply with the word "radius" and I'll send you the audit sheet.

🧰 Trending Tools

  • XBOW — the autonomous offensive-security agent that found the Bing RCEs. For teams that own real production surface.

  • Aikido Security — AI pentest agents that reviewed NodeBB's source in six hours. For anyone shipping a codebase without a security hire.

  • Okta for AI Agents — governs agent connections at runtime, not just at setup. For anyone whose agents hold credentials.

  • Codenotary AgentMon 3 — learns your agents' normal behavior and adapts policy when they deviate. For small teams running agents unattended.

  • Jamf AI Governance — discovers which AI tools your people actually use and produces audit-ready reporting. For anyone who suspects the answer is "more than I approved."

📣 Put your brand here. The AI Innovator reaches AI-first operators, creators, and marketers every weekday. Primary sponsorships are now booking.

That's a wrap

Tomorrow: whether anyone reproduces the Redis findings with logs attached — the claim that decides how seriously to take agentic bug hunting.

Enjoying the format? Share your link — every referral keeps this free:

How was today’s email?

(Tell us what you liked or what could be better)

Login or Subscribe to participate

You're reading the 3-minute AI Innovator — same time, every weekday. Hit reply and tell me what you want more of. I read every one.

PARTNER SPOTLIGHT

ListGuru — ask for anyone, get their verified contact

Describe who you need in plain English. ListGuru searches public profiles across LinkedIn, X, GitHub and the open web — then reveals verified emails and mobile numbers you can actually reach.

  • No boolean, no filter matrices — type "Heads of Growth at Series-B fintechs hiring SDRs" and its agents research, qualify and rank the matches.

  • Verified at reveal, never cached — syntax, disposable-domain and role-account checks plus a live mail-server lookup the moment you unlock a contact.

  • You pay for results, not attempts — a search that finds nobody costs zero credits.

  • 25 free credits every month, no card required. Paid plans start at $29/mo.