TLDRocket
Sign in

Apple and Bynario agree GPT-5.5 found a real macOS bug. They disagree on the report cap.

The New Stack Amanda Caswell

Apple caps open security reports, and hitting it means a 30-day wait to file another bug. An AI still found a real macOS root-privilege flaw anyway — it just had to wait in line.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Apple quietly tightened its security bug bounty rules back in June. Researchers can now only have so many reports open at once, and once they hit that ceiling, they're stuck waiting 30 days before they can flag anything else through the company's internal portal. The reason is blunt: Apple got buried in AI-generated submissions, and a lot of them turned out to be nothing — junk flaws that looked plausible on paper but didn't hold up once someone actually tried to reproduce them.

Then Bynario, an Italian cybersecurity firm, ran into the new limit the hard way. Using GPT-5.5 through its Atlas platform, the company surfaced more than 50 possible bugs in Apple's latest Mac operating system in just three weeks. One of them was real: a flaw in macOS Screen Sharing that let an authenticated VNC user reach protected data and create files with root privileges. Apple assigned it CVE-2026-43760 and shipped a fix in macOS Tahoe 26.6. But Bynario couldn't even report it right away, because it had already maxed out its open-investigation slots. Apple has since reached out to go through Bynario's backlog directly.

Apple isn't just on the receiving end of this shift. It's using AI internally to hunt vulnerabilities in its own code, and its recent updates carried roughly five times the security fixes of earlier release cycles. Across the industry, similar tools are compressing work that used to take weeks — prompt injection testing, agent hardening — into minutes. HackerOne has built something comparable called Hai Triage, which reviews incoming reports in minutes and sorts them as likely valid, invalid, or needing a closer look, complete with a suggested severity and a recommendation on whether a human analyst should bother. The company insists humans still make every final call.

That's the real tension here. A submission cap doesn't know the difference between a convincing pile of AI slop and a genuine flaw an AI happened to find first. Apple's existing bounty rules already require affected versions, reproduction steps, and a proof of concept — and researchers who repeatedly send in unvalidated theoretical bugs risk a 180-day pause, with permanent removal after two strikes. One fix floated is letting researchers with a proven track record submit more freely while asking AI-heavy submitters to demonstrate reproducibility up front. Without something like that, legitimate findings sit in the same queue as noise.

Apple's triage problems predate this AI wave, too. Tyler Murphy, co-founder of EasyOptOuts, first flagged a flaw in Apple's Hide My Email feature back in June 2025 — one that could expose the real addresses the feature was supposed to hide. It took a year of back-and-forth before he went public. Apple said it patched the issue on July 3, 2026, but AppleInsider reproduced the same behavior two weeks later, on July 17. EasyOptOuts now says it's fixed, though addresses exposed before the working patch may still be sitting in email records somewhere.

Apple's AI ambitions aren't slowing down either — the company recently turned Safari into something AI agents can operate directly, which only adds more surface area that eventually needs the same overworked triage pipeline. The bug volume problem and the product expansion problem are pointing in the same direction at the same time.

My take — AI-written commentary, not fact-checked reporting

Apple didn't build this cap because AI is bad at finding bugs — Bynario just proved the opposite by turning up a real root-privilege flaw in three weeks. The actual bottleneck is that humans still have to verify every single report, and no amount of clever AI triage on Apple's side changes that math once submission volume explodes. A flat 30-day timeout punishes good and bad reports identically, which is exactly the kind of blunt fix you get when a company is drowning and reaching for the nearest lever. If Apple can lean on the same kind of sorting logic HackerOne is already running, that seems like a far better bet than making a researcher who just found a legitimate root-privilege exploit sit in a queue behind AI slop.”}]}

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.