Interconnects
·
3 months ago
Anthropic announced Claude Mythos, a cybersecurity-focused model with strong stated abilities, prompting renewed warnings that open-weight versions would enable widespread attacks on unprepared infrastructure. Running such a model requires approximately 100 H100 GPUs costing roughly $10,000 daily, meaning only well-resourced actors could deploy it rather than individual attackers. The author argues for specific research into the actual cybersecurity risks and capability gaps between open and closed models rather than broad restrictions that could cede AI development to other countries.
Deep Learning Weekly
·
3 months ago
Deep Learning Weekly Issue 450 covers major releases including Google's Gemma 4 open models, Alibaba's Qwen3.6-Plus agentic coding model, and Anthropic's Claude Managed Agents in public beta, alongside papers on video object removal and KV cache compression for long-context reasoning. Google's Gemma 4 31B model ranks third among open models on Arena AI, and Ollama 0.19 delivers approximately 2x speed gains on Apple Silicon M5 chips using MLX-powered inference. The releases enable faster production deployment of AI agents and more efficient inference for extended reasoning tasks.
Google Research
·
3 months ago
Researchers introduced ConvApparel, a dataset of over 4,000 human-AI conversations, to measure and reduce the realism gap in LLM-based user simulators used for testing conversational AI agents. The dataset comprises nearly 15,000 turns collected through a dual-agent protocol where participants interacted with either helpful or intentionally unhelpful shopping assistants, with fine-grained turn-by-turn annotations of user satisfaction and frustration. Data-driven simulators (in-context learning and supervised fine-tuning) demonstrated superior performance and realistic adaptation to novel scenarios compared to prompt-based approaches, suggesting that improved training methods can better prepare conversational agents for real-world deployment.
Hugging Face Blog
·
3 months ago
Sentence Transformers library version 5.4 now enables users to encode and compare text, images, audio, and video using a unified API for embedding and reranking tasks. The Qwen3-VL-Embedding-2B model requires approximately 8 GB of GPU VRAM, with 20 GB needed for 8B variants. This multimodal capability allows new applications such as visual document retrieval, cross-modal search, and retrieval-augmented generation pipelines that combine different input types.
Hugging Face Blog
·
3 months ago
Overworld released Waypoint-1.5, a real-time video world model designed to run on consumer GPUs rather than datacenter clusters. The 720p version runs at 60 FPS on RTX 3090-5090 hardware, while a 360p tier supports gaming laptops and future Apple Silicon Macs, trained on nearly 100 times more data than the original Waypoint. Users can now run interactive generative worlds locally through the Biome desktop client or instantly via browser on Overworld Stream without datacenter resources.
OpenAI Blog
·
3 months ago
● 3 sources
CyberAgent adopted ChatGPT Enterprise and Codex to accelerate operations across its advertising, media, and gaming divisions while maintaining security standards. The company uses these tools to improve work quality and speed up decision-making processes across multiple business units. This shift allows CyberAgent to scale AI use across the organization without compromising data protection requirements.