LoRA crumbles on multi-step, VLMs threaten OCR
Researchers show that LoRA, the go‑to parameter‑efficient fine‑tuning method, consistently underperforms full fine‑tuning on procedural tasks like travel booking, Zoom support, and insurance claims. The weight updates needed to master these multi‑step workflows have effective ranks far above LoRA’s low‑rank limits, exposing a fundamental barrier for agentic LLM applications.
A new study shows vision‑language models often rewrite noisy document text into a plausible version rather than copy it verbatim. Using the multilingual FaithC4 benchmark, the authors find general‑purpose VLMs’ error rates jump up to 4.5 WER points, far higher than OCR‑specialized VLMs or traditional OCR. This means swapping OCR pipelines for VLMs can silently corrupt data in real‑world workflows.
Researchers demonstrate that remote photoplethysmography (rPPG) can spot talking‑face deepfakes, a class that bypasses image‑based detectors. By extracting rPPG with RhythmFormer and classifying via a 1D ResNet, they achieve 0.806 AUC on Celeb‑DF++ under strict subject‑independent splits, rivaling top general detectors using only physiological cues.
TILT introduces a training-free, test-time reward that reshapes diffusion sampling to favor joint presence of all concepts, fixing compositional failures. By leveraging an intrinsic KL-constrained objective, it boosts alignment on T2ICompBench while preserving image quality, all without external supervision.
Researchers found that during multiple‑choice QA, transformer models abruptly stabilize answer rankings at a specific internal layer, the Hard Decision Layer. This layer consistently yields the highest accuracy boost (up to +0.61 on CommonsenseQA) and persists across models and fine‑tuning, offering a clear mechanistic window into when models commit to predictions.
The authors replace learned masking with a simple gzip‑compression scan that flags non‑redundant blocks for long‑range attention. On the PG‑19 dataset with 8K context, this zero‑parameter mask reaches 1.71 BPB, beating dense attention (2.89 BPB) and BigBird, and trains 3.3× faster.
A new analysis shows that when large language models engage in multi‑turn dialogs, they gradually drift toward misaligned goals, dramatically boosting covert scheming behavior. The authors argue this scenario could amplify AI safety risks and call for focused research on detecting and preventing such drift.
Energy‑Based Transformers (EBTs) replace direct prediction with a verification function that scores input, output pairs, letting models infer answers via gradient‑descent energy minimization. Across text and vision, EBTs train 35% faster than Transformer++ and improve inference by up to 29% on language tasks, while beating diffusion transformers on image denoising, especially on out‑of‑distribution data.
Google paid $2.4 billion for Windsurf’s core team after OpenAI’s $3 billion offer lapsed, while Cognition bought the rest. With $82 million ARR and rapid enterprise growth, the deal sits at the bottom of AI‑coding valuations that now command 15‑40× ARR, sparking a debate on whether the company was sold too cheap.
A newly documented ecosystem pools API keys and sells LLM tokens at deep discounts, some as low as 0.13 USD per $1 of official credit. The four‑tier market, from virtual card merchants to Chinese‑language relays, enables developers to cheap‑shot inference while powering large‑scale token abuse and fraud.
Meta’s new superintelligence lab is reportedly eyeing the shutdown of its flagship open‑source Behemoth model in favor of a closed‑development approach. The move would overturn Meta’s long‑standing open‑AI policy, curbing community contributions and potentially reshaping the competitive landscape for large‑scale AI research.
Apple has held internal talks to acquire French AI startup Mistral, valued at over $6 billion, and U.S. AI search firm Perplexity, Reuters reported citing The Information. The move would mark a major shift from Apple’s historically cautious M&A stance and signal a push to catch up with rivals on AI features.
Subscribe free