LodeHQSubscribe →

Mage 4B, Gemini 3.6 Flash, AI matches chemists

AI · 2026-07-23

Models & Releases
Microsoft launches Mage: 4‑billion‑parameter multimodal models that rival 20B‑plus systems3 MIN

Mage delivers a 4 B‑parameter vision‑language model and a text‑to‑image diffusion stack that match or beat much larger open models while running on a single A100. Its compact VAE and native‑resolution diffusion cut compute and speed up training, and checkpoints are now on Hugging Face.

Google rolls out Gemini 3.6 Flash, boosting token efficiency for AI agents5 MIN

Google announced three Gemini variants: 3.6 Flash, which cuts output token use by 17% and halves agentic task costs; 3.5 Flash‑Lite, a 350‑token‑per‑second low‑latency model; and 3.5 Flash‑Cyber, a security‑focused model paired with CodeMender. The launches promise cheaper, faster production AI agents and tighter cyber‑security integration.

Research
Google Quantum AI trains RL agent to continuously correct qubit drift5 MIN

They demonstrated a reinforcement‑learning framework that learns from quantum error detections to continuously steer control parameters, enabling error correction without halting computations. This removes a bottleneck for long‑duration quantum algorithms and pushes fault‑tolerant quantum computing closer.

Evaluation Pipelines Skew Sparse Autoencoder Interpretability Scores2 MIN

Systematic tests show that autointerpretability scores for sparse autoencoders vary more with pipeline choices, prompt templates and scoring language models, than with the underlying model architecture. This undermines cross‑paper comparisons and calls for standardized evaluation protocols, including a variance decomposition tool and a reporting checklist.

CrackedPDFs exposes hidden prompt injection risk in PDF‑based LLM pipelines1 MIN

A new benchmark called CrackedPDFs shows that LLM systems that flatten PDFs can miss hidden instructions injected into the document, creating a prompt‑injection vector for Retrieval‑Augmented Generation. The dataset includes over 29k PDFs and the hybrid detector reaches 0.96 F1, proving that structural‑text hybrid defenses can catch many hidden attacks.

Agentic AI Matches Grad Chemists in NMR Structure Elucidation1 MIN

The paper shows an autonomous AI agent, guided by a frozen LLM, that treats NMR interpretation as a constrained search rather than direct prediction. On benchmark datasets it reaches 71% top‑1 accuracy, on par with graduate students (66%) and beating zero‑shot deep models. This reframes spectroscopic analysis toward tool‑orchestrated reasoning.

Policy & Safety
NEXUS adds real-time safety checks to tool-using LLM agents1 MIN

NEXUS is a structured-plan monitor that evaluates LLM agent tool calls and can allow, block, request human confirmation, or sandbox the action. On synthetic benchmarks it hits 0.949 F1 and keeps latency under 0.3 ms, offering low-overhead runtime safety for high-impact agents.

Tools & Open Source
Nunchaku 4‑bit Quantization Slashes Diffusers Memory and Latency10 MIN

Diffusers now supports Nunchaku’s 4‑bit SVDQuant inference, letting you run diffusion models with W4A4 weights and activations. The change halves GPU memory and cuts generation time to under 2 seconds for 1024×1024 images on a RTX 5090, removing the need for custom kernels.

ACP v2 draft released, inviting community feedback on richer editor-agent workflows4 MIN

The Agent Client Protocol v2 draft is now public. It expands the spec to allow background updates, streaming tool calls, and granular diff structures, letting editors and AI coding agents coordinate beyond single user turns. The team is seeking feedback to refine these new session and permission features.

NVIDIA open-sources GPU-accelerated framework for realistic medical-physics simulation3 MIN

NVIDIA released an open-source, GPU-accelerated Medical Physics Simulation framework within its Isaac for Healthcare platform. It lets developers model tissue-instrument interaction, generate hard-to-capture edge cases, and run thousands of parallel simulations, slashing training time from hours to minutes. The transparency speeds robot validation and aids regulatory review.

Get AI in your inbox, every issue.
Subscribe free
Get the app · Privacy · Terms · About · Contact
© 2026 LodeHQ