TLDRocket
Sign in

Launch HN: General Instinct (YC P26) – Frontier models on edge devices

Hacker News guanming0717

General Instinct, a YC-backed startup, released InstinctRazor, a tool for compressing large language models to run on edge devices with limited compute. They compressed Qwen3.5-122B from 245 GB to 48 GB while matching or exceeding the performance of smaller models like Gemma-4-26B on benchmarks. The technique enables frontier models to run on robotics and edge hardware with 7.6–8 GB peak VRAM usage, addressing the gap between datacenter-optimized models and resource-constrained physical systems.

Why it matters

Hey HN, Guanming and Bill here from General Instinct (https://general-instinct.com/).After years of working in robotics, we kept running into the same problem: the best models never fit the hardware we actually had available.The models that performed best were usually designed around datacenter assumptions: large GPUs, lots of memory bandwidth, and reliable network access. But most physical systems have the opposite constraints.That led us down the path of figuring out how much of a frontier model could be preserved while still making it practical to run on edge hardware.As part of that work, we recently open sourced InstinctRazor (https://github.com/General-Instinct/InstinctRazor)One result we're excited about is compressing Qwen3.5-122B-A10B, a roughly 245 GB BF16 MoE model, into a 48 GiB GGUF. The resulting model is actually smaller than Gemma-4-26B-A4B while outperforming it on benchmarks like MMLU-Pro and GPQA-D etc. we preserve the parts that are always active (router, norms, Gat

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.