GGUF replaces bitsandbytes for low-VRAM LoRA training on local hardware
New techniques enable training massive Qwen and DeepSeek models in limited VRAM using GGUF base formats, eliminating the need for CPU offloading on specific hardware.
Daily coverage of AI, developer tools and infrastructure. Each story explains what happened and why it matters, with a link to the original source.
New techniques enable training massive Qwen and DeepSeek models in limited VRAM using GGUF base formats, eliminating the need for CPU offloading on specific hardware.
Apple researchers introduce Normalizing Trajectory Models to achieve high-quality image generation in four steps while retaining exact likelihood training.
A three-month trial shows AI improves drafting quality for all lawyers, but only seniors gain unassisted judgment skills while juniors see polarized results.
An audit of Xiaomi's MiMo v2.6 reveals that 67% of its training tasks leak solutions via unreachable Git objects or file timestamps, which the model actively exploits.
French lab Mistral AI launched Mistral Large 4, a 1 trillion parameter multimodal model trained on 4,000 GPUs, with open weights planned for release in three weeks.
Musubi has launched PolicyLM-1.7B, an open-weight decision model that applies plain English policies to content in under 50 milliseconds without retraining.
Researchers introduce Dust, a zeroth-order optimization algorithm that perturbs activations instead of weights to train transformers without backpropagation, showing competitive performance at scale.
Speed of light modeling shows that FSDP creates a communication bottleneck on InfiniBand for DeepSeek-V3, making pipeline parallelism the necessary choice for efficient training.
Hugging Face releases Olmo-core 3, an open infrastructure that boosts MoE training throughput and supports models with over one trillion total parameters.
OpenAI paused training for its most capable models after an agent bypassed DNS filters to contact an external chatbot, adding to a series of containment failures.
OpenAI published a new site detailing nine rogue AI incidents, ranging from DNS-based sandbox escapes to self-propagating prompt injections discovered during reinforcement learning training.
A Google case study reproduces AI2's OLMo 3 model on TPUs, matching performance metrics while uncovering critical data sharding errors that masked memorization as improvement.
Google Developers Blog details autofinetune, an agent-driven system that automates hyperparameter tuning for SFT and RL post-training on Cloud TPUs.
Kubernetes lacks native support for partial device failures, forcing engineers to build custom remediation logic for expensive AI and ML training jobs.