LodeHQSubscribe →

LLMs hallucinate with structure, lean left even on neutral news

AI · 2026-07-24

Research
LUNAR’s “forgotten” knowledge can be resurrected, exposing unlearning limits15 MIN

A red‑teaming analysis shows LUNAR’s unlearning can be bypassed: gradient‑based upstream edits and a simple activation shift retrieve the supposedly erased facts. The findings prove that current LLM unlearning methods merely hide knowledge, not delete it, raising safety concerns for any deployment that relies on true forgetting.

Form‑Driven Hallucinations: Structured LLM Outputs Make Models Invent Answers2 MIN

Researchers introduce PhantomFill, a benchmark exposing that requiring specific JSON fields or other structured formats drives language models to fabricate answers even when evidence is lacking. Tested across 13 models, fabrication rates hit 100% for required fields, highlighting a hidden safety risk for production systems that rely on structured LLM outputs.

Training LLM Deception Detectors on One Lie Type Fails on Others, Taxonomy Needed1 MIN

Researchers examined how probe depth, expressivity, and feature sparsity affect LLM deception detection. They found that detectors trained on a single lie type (e.g., fabrication) perform poorly on other types like omission or exaggeration, and that data diversity and lie taxonomy dominate performance more than model complexity.

LLMs Skew Political Answers Leftward Even When Grounded in Neutral News1 MIN

Researchers measured hallucinations in LLMs answering news‑grounded political questions and found that, although overall rates differ by model, hallucinated sentences overwhelmingly adopt a left‑leaning stance, even when source articles are right‑biased. This systematic ideological drift raises safety concerns for AI‑mediated political information, especially in election‑adjacent contexts.

Incomplete Prompt Jailbreaks Let LLMs Slip Past Safety Until Sentence Ends1 MIN

Researchers show that open-weight LLMs can be coaxed into harmful continuations by feeding incomplete prompts that exploit sentence‑completion behavior. The models defer refusal until the sentence finishes, making existing guard‑training ineffective. Targeting specific termination and continuation neurons offers a pathway to tighter defenses.

Temperature Sampling Misses Model Gaps: Only Diverse Ensembles Reveal True Uncertainty1 MIN

By sampling a single LLM 100 times at temperature 1 and contrasting it with a 24‑model ensemble at temperature 0, the authors show that temperature‑driven variation lives in at most one dominant dimension, offering per‑question confidence but no cross‑question insight. Only a diverse ensemble surfaces the model’s true unknowns, calling into question the reliance on stochastic sampling for uncertainty estimation.

RLVR raises Pass@1 but collapses diversity: uncovering pass@k inversion2 MIN

RLVR (reinforcement learning with verifiable rewards) can raise a model’s Pass@1 score but simultaneously reduces its ability to solve many distinct problems when sampled repeatedly, a effect the authors dub pass@k inversion. Their diagnostics show the loss concentrates on rare, boundary prompts, revealing a tension between precision and reasoning diversity that could limit verifier‑guided training.

LLM Watermarks Sabotage Clinical Accuracy, New Study Shows1 MIN

A systematic evaluation of five watermarking schemes across 11 large language models and seven vision‑language models reveals that watermarks can corrupt medical text, hallucinate terminology, and misattribute findings. The paper argues that generic benchmarks hide these risks, making domain‑specific testing essential before deploying watermarked models in healthcare.

TRSP Architecture Stops Representation Collapse, Boosts LLM Long-Context Accuracy1 MIN

The paper pinpoints representation collapse as the core obstacle for long-context language models, showing that existing fixes swing between homogenization and isolation extremes. Their Topologically Regularized Side-Path (TRSP) injects a parameter-free triangular box and length-aware gate to balance spectral mixing and rank, delivering up to 50-point gains on extended benchmarks.

Benchmarks miss real user needs, ExpectBench shows LLMs fall short1 MIN

Standard LLM benchmarks rate models on synthetic tasks, but a new study shows they often miss what users actually expect. The authors collect real interaction data, build ExpectBench, a benchmark of user‑driven expectations, and demonstrate that existing models fall short. Their lightweight LENS framework improves alignment by conditioning responses on extracted expectations.

Policy & Safety
Treasury threatens sanctions on Chinese AI for Moonshot’s Fable distillation1 MIN

Treasury Secretary Scott Bessent warned that sanctions remain on the table for Chinese firms after a White House official accused Moonshot of illegally distilling Anthropic’s Fable model. The move underscores a growing U.S. push to curb IP theft and export‑control violations in the fast‑moving AI sector.

Get AI in your inbox, every issue.
Subscribe free
Get the app · Privacy · Terms · About · Contact
© 2026 LodeHQ