TLDRocket
Sign in

Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary

MarkTechPost Michal Sutter

Pokee AI released Isaac, a 28B model that remembers 10M tokens and runs on-prem. That's huge for banks, hospitals, and gov work that can't touch cloud AI.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Agents that run for a long time have a problem nobody talks about enough: they pile up context faster than they finish anything. Every tool call, every observation, every half-formed thought stays in the window, and until now the only systems that could actually hold onto that much material and stay coherent while doing it lived behind a cloud API. Fine for a startup. Not fine for a hospital system, a defense contractor, or anyone else whose data isn't allowed to leave the building. Pokee AI built Pokee-Isaac 28B specifically for that gap — a 28-billion-parameter, text-only model with a 10-million-token context window, meant to be deployed inside the customer's own walls rather than rented from someone else's.

The long-context numbers back up the pitch. Isaac holds 93.3% on RULER all the way out to 10M tokens, and according to Pokee's own panel, every baseline model it tested against drops to 0.0 past 2M. GPT-5.6 Luna and Gemini 3.5 Flash Lite keep pace up to 512K, then simply overflow at the 1M mark. On MRCR v2 with eight needles buried in the text, Isaac scores 0.607, 0.743, and 0.500 at 256K, 512K, and 1M respectively, and its lead over Gemini actually grows as the context gets longer, from a 0.133 gap to 0.295.

On the agentic side, the story is closer to a tie than a rout, and Pokee is upfront about that. Isaac edges BFCL v4 at 70.94 against Luna's 70.61 — parity, not domination, and the report says so directly. On τ³-bench it averages 0.662 across four domains versus Gemini's 0.631, though the banking domain is brutal for everyone at 0.186. On MCP-Atlas it lands third at 74.59% coverage but gets there using only 9.10 turns per task, compared to Gemini's 14.99 — fewer steps, less compute burned per task. Terminal-Bench 2.1 is the one place a cloud model wins outright: Isaac clears 56 of 86 tasks (65.1%) against Luna's 60, and the report doesn't try to spin it. On the security side, DTAP red-teaming puts Isaac at the lowest attack success rates across direct, indirect, and combined categories (36.0, 35.2, 35.6) while keeping 82.5% benign task success — though it's worth flagging that Isaac ran under Pokee's own harness while the baselines used the stock runner.

The efficiency numbers matter as much as the accuracy ones if you're actually planning to host this thing. On a single B200-class GPU, time-to-first-token runs 23.6 seconds at 1M tokens and 72.9 seconds at 10M — a tenfold jump in context for roughly a threefold hit to latency. Prefill throughput actually climbs with longer context, from 42,400 to 137,200 tokens per second. Pricing is listed provisionally at $0.15 per million input tokens and $1.00 per million output tokens. Pokee says single-GPU serving is possible starting from something like an RTX 4090, though that specific claim is vendor guidance rather than a measured result, since the published numbers all come from B200-class hardware. On-device support extends to Intel's Arc Pro B70 and Core Ultra Series 3 chips, plus Qualcomm's Snapdragon X2 Elite.

None of this is open-weight. Isaac ships through an OpenAI-compatible API or gets licensed for VPC, on-prem, or on-device deployment, with day-zero support for vLLM and SGLang. That's a deliberate choice aimed at a specific customer: mid-size and enterprise teams that already run their own inference stack, plus device makers — healthcare, finance, defense, legal, and pharma being the obvious fits, since the common thread across all of them is a hard rule against data leaving the building, not just a preference for privacy.

My take — AI-written commentary, not fact-checked reporting

The licensing choice here is more interesting than the leaderboard numbers. Pokee isn't trying to win an open-weights popularity contest; it's betting that regulated industries will pay for a hard boundary rather than wait for someone to open-source something they could self-host anyway. That's a shrewder read of who actually needs 10 million tokens of memory than chasing another benchmark trophy. The tradeoff is real, though: locking a hospital or a defense contractor into a closed license for their long-context agent means they're trusting Pokee's terms and pricing to stay reasonable long after the launch blog post fades.Wait, the instructions say opinion shouldn't use

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.