AutoJudge accelerates large language model inference by automatically identifying which mismatched tokens between draft and target models don't affect task correctness, eliminating the need for manual annotation. The method achieves 1.5–2x speedups over standard speculative decoding while accepting up to 40 draft tokens per verification cycle with minimal accuracy loss, and integrates into existing frameworks like vLLM and TensorRT-LLM. Inference speed improves across benchmarks with only 2–4% accuracy drops: on GSM8K, the Llama-3.1-70B/8B pair reaches 107.4 tokens/s (1.49x faster), and on code tasks, acceptance rates increase 2.3–3.5x.
TorchForge reinforcement learning pipelines now run on Together AI's Instant Clusters with support for distributed training across GPU and CPU nodes. The demo trains a Qwen 1.5B model to play BlackJack using GRPO through a pipeline integrating vLLM, Monarch, and TorchTitan, deployable with three kubectl commands. This infrastructure enables RL agents to tackle diverse tasks from game-playing to coding through unified pipeline architecture with sandboxed environments.
Together AI and Meta partnered to integrate PyTorch Reinforcement Learning capabilities into Together's AI cloud platform, enabling users to build and deploy reinforcement learning agents. The integration provides access to Meta's open-source PyTorch RL tools directly within Together's infrastructure. This allows developers to train and deploy reinforcement learning models more easily on Together's cloud platform.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.