Learn how Cursor partnered with Together AI to deliver real-time, low-latency inference at scale
Together AI
Cursor partnered with Together AI to optimize real-time code completion inference using NVIDIA Blackwell hardware and advanced quantization techniques. The partnership productionized B200/GB200 chips with FP4 quantization and TensorRT optimizations to achieve low-latency inference at scale. This infrastructure enables Cursor's in-editor AI agents to maintain fast response times while processing requests reliably.
Why it matters
Together AI teamed with Cursor to build the real-time inference stack that keeps in-editor agents fast and reliable. They productionized NVIDIA Blackwell (B200/GB200), tuning ARM hosts, kernels, and FP4/TensorRT quantization for low latency and rapid model rollouts.