Google Gemini 1.5 Flash-8B drops API pricing to $0.0375 per million tokens
Google announced broad availability of Flash-8B with sub-100ms response times, designed for ultra-high-volume embedding, summarization, and agent classification tasks.
Autonomous refactoring agents, open-source video synthesizers, and real-time audio models lead early autumn announcements.
Google announced broad availability of Flash-8B with sub-100ms response times, designed for ultra-high-volume embedding, summarization, and agent classification tasks.
LiveKit open-sourced their end-to-end voice agent framework connecting WebRTC, Whisper, LLMs, and ElevenLabs with sub-400ms glass-to-glass latency.
THUDM released CogVideoX-5B with permissive licensing, capable of generating 1080p 60-frame videos on a consumer RTX 3090/4090 with FP8 quantization.
Codestral received a major checkpoint with Fill-In-the-Middle (FIM) across 256K tokens, enhancing inline tab-completion across giant multi-repo codebases.
Supabase integrated automated vector generation directly inside Postgres triggers, executing embeddings locally without external Node/Python microservices.
Hugging Face released SmolLM2, outperforming previous 3B models on common sense and code syntax while consuming under 1.2GB RAM on mobile phones.
OpenAI released the Turbo variant of Whisper Large v3, reducing decoder layers from 32 to 4 with negligible accuracy degradation across 90+ languages.
Cerebras launched commercial API endpoints demonstrating instantaneous generation speeds exceeding human reading limits by 50x.
Anthropic published internal evaluation test suites that grade prompt edge cases against regression tests before deploying updates to production.
European regulators issued technical compliance templates for general-purpose AI models, requiring documented training data transparency and copyright filters.
4 additional fast-moving updates and developer breakthroughs
Serve 50+ personalized user fine-tunes concurrently with sub-millisecond adapter switching overhead.
Zero-ad, index-level raw search extraction providing an alternative to Google & Bing APIs for RAG agents.
Enables seamless graph captures of variable-length attention masks with 20% lower compilation latency.
Developers can now inject bezier camera trajectory curves directly via JSON into Gen-3 generation calls.
New editions are published every Sunday morning. Subscribe to our free newsletter or feed to never miss a breakthrough.