Mage 4B, Gemini 3.6 Flash, AI matches chemists
Mage delivers a 4 B‑parameter vision‑language model and a text‑to‑image diffusion stack that match or beat much larger open models while running on a single A100. Its compact VAE and native‑resolution diffusion cut compute and speed up training, and checkpoints are now on Hugging Face.
Google announced three Gemini variants: 3.6 Flash, which cuts output token use by 17% and halves agentic task costs; 3.5 Flash‑Lite, a 350‑token‑per‑second low‑latency model; and 3.5 Flash‑Cyber, a security‑focused model paired with CodeMender. The launches promise cheaper, faster production AI agents and tighter cyber‑security integration.
They demonstrated a reinforcement‑learning framework that learns from quantum error detections to continuously steer control parameters, enabling error correction without halting computations. This removes a bottleneck for long‑duration quantum algorithms and pushes fault‑tolerant quantum computing closer.
Systematic tests show that autointerpretability scores for sparse autoencoders vary more with pipeline choices, prompt templates and scoring language models, than with the underlying model architecture. This undermines cross‑paper comparisons and calls for standardized evaluation protocols, including a variance decomposition tool and a reporting checklist.
A new benchmark called CrackedPDFs shows that LLM systems that flatten PDFs can miss hidden instructions injected into the document, creating a prompt‑injection vector for Retrieval‑Augmented Generation. The dataset includes over 29k PDFs and the hybrid detector reaches 0.96 F1, proving that structural‑text hybrid defenses can catch many hidden attacks.
The paper shows an autonomous AI agent, guided by a frozen LLM, that treats NMR interpretation as a constrained search rather than direct prediction. On benchmark datasets it reaches 71% top‑1 accuracy, on par with graduate students (66%) and beating zero‑shot deep models. This reframes spectroscopic analysis toward tool‑orchestrated reasoning.
NEXUS is a structured-plan monitor that evaluates LLM agent tool calls and can allow, block, request human confirmation, or sandbox the action. On synthetic benchmarks it hits 0.949 F1 and keeps latency under 0.3 ms, offering low-overhead runtime safety for high-impact agents.
Diffusers now supports Nunchaku’s 4‑bit SVDQuant inference, letting you run diffusion models with W4A4 weights and activations. The change halves GPU memory and cuts generation time to under 2 seconds for 1024×1024 images on a RTX 5090, removing the need for custom kernels.
The Agent Client Protocol v2 draft is now public. It expands the spec to allow background updates, streaming tool calls, and granular diff structures, letting editors and AI coding agents coordinate beyond single user turns. The team is seeking feedback to refine these new session and permission features.
NVIDIA released an open-source, GPU-accelerated Medical Physics Simulation framework within its Isaac for Healthcare platform. It lets developers model tissue-instrument interaction, generate hard-to-capture edge cases, and run thousands of parallel simulations, slashing training time from hours to minutes. The transparency speeds robot validation and aids regulatory review.
Subscribe free