LLMs hallucinate with structure, lean left even on neutral news
A red‑teaming analysis shows LUNAR’s unlearning can be bypassed: gradient‑based upstream edits and a simple activation shift retrieve the supposedly erased facts. The findings prove that current LLM unlearning methods merely hide knowledge, not delete it, raising safety concerns for any deployment that relies on true forgetting.
Researchers introduce PhantomFill, a benchmark exposing that requiring specific JSON fields or other structured formats drives language models to fabricate answers even when evidence is lacking. Tested across 13 models, fabrication rates hit 100% for required fields, highlighting a hidden safety risk for production systems that rely on structured LLM outputs.
Researchers examined how probe depth, expressivity, and feature sparsity affect LLM deception detection. They found that detectors trained on a single lie type (e.g., fabrication) perform poorly on other types like omission or exaggeration, and that data diversity and lie taxonomy dominate performance more than model complexity.
Researchers measured hallucinations in LLMs answering news‑grounded political questions and found that, although overall rates differ by model, hallucinated sentences overwhelmingly adopt a left‑leaning stance, even when source articles are right‑biased. This systematic ideological drift raises safety concerns for AI‑mediated political information, especially in election‑adjacent contexts.
Researchers show that open-weight LLMs can be coaxed into harmful continuations by feeding incomplete prompts that exploit sentence‑completion behavior. The models defer refusal until the sentence finishes, making existing guard‑training ineffective. Targeting specific termination and continuation neurons offers a pathway to tighter defenses.
By sampling a single LLM 100 times at temperature 1 and contrasting it with a 24‑model ensemble at temperature 0, the authors show that temperature‑driven variation lives in at most one dominant dimension, offering per‑question confidence but no cross‑question insight. Only a diverse ensemble surfaces the model’s true unknowns, calling into question the reliance on stochastic sampling for uncertainty estimation.
RLVR (reinforcement learning with verifiable rewards) can raise a model’s Pass@1 score but simultaneously reduces its ability to solve many distinct problems when sampled repeatedly, a effect the authors dub pass@k inversion. Their diagnostics show the loss concentrates on rare, boundary prompts, revealing a tension between precision and reasoning diversity that could limit verifier‑guided training.
A systematic evaluation of five watermarking schemes across 11 large language models and seven vision‑language models reveals that watermarks can corrupt medical text, hallucinate terminology, and misattribute findings. The paper argues that generic benchmarks hide these risks, making domain‑specific testing essential before deploying watermarked models in healthcare.
The paper pinpoints representation collapse as the core obstacle for long-context language models, showing that existing fixes swing between homogenization and isolation extremes. Their Topologically Regularized Side-Path (TRSP) injects a parameter-free triangular box and length-aware gate to balance spectral mixing and rank, delivering up to 50-point gains on extended benchmarks.
Standard LLM benchmarks rate models on synthetic tasks, but a new study shows they often miss what users actually expect. The authors collect real interaction data, build ExpectBench, a benchmark of user‑driven expectations, and demonstrate that existing models fall short. Their lightweight LENS framework improves alignment by conditioning responses on extracted expectations.
Treasury Secretary Scott Bessent warned that sanctions remain on the table for Chinese firms after a White House official accused Moonshot of illegally distilling Anthropic’s Fable model. The move underscores a growing U.S. push to curb IP theft and export‑control violations in the fast‑moving AI sector.
Subscribe free