AI startups 25% smaller; evil vector is dread
The post argues that steering vectors, inoculation prompting, and post‑hoc honesty fine‑tuning are just variants of a single “train‑deploy mismatch” framework, training a model under one configuration and deploying it under another. This shared structure creates the same trade‑off between training relevance and deployment efficacy, shaping how we should think about future alignment work.
Researchers trained a neologism token on model‑generated data and found the model’s self‑labeled “evil” steering vector aligns with the emotion of dread rather than malicious intent. This mismatch shows that asking models to self‑report their internal concepts can be misleading, raising fresh concerns for interpretability work.
A new Harvard Business School/INSEAD working paper finds AI‑native firms are about a quarter smaller and have flatter hierarchies, with more engineers and fewer managers. Despite leaner staff, they achieve higher per‑employee valuations, showing product‑embedded AI can scale knowledge work without expanding headcount.
Using Wilsonian renormalization group theory, the authors treat attention as a perturbation of the MLP residual stack. They find attention is a relevant operator for long‑range data, driving a phase transition and higher effective rank, but irrelevant for short‑range data, where the Transformer behaves like the MLP. The work links attention’s utility to data spectral structure.
The paper introduces Recursive Harness Self‑Improvement (RHI), a prompt‑level loop that refines user‑built harnesses with pairwise feedback. A few RHI cycles let low‑reasoning agents surpass high‑reasoning baselines while slashing inference cost up to 60%, driven by better task context flow.
A test on 4,181 Omni‑MATH problems using gpt‑oss‑120b agents shows that, despite higher precision, reviewer stages rarely change the next answer, especially on harder tiers. Broadcast‑style peer discussion outperforms the planner‑executor‑reviewer pipeline, exposing a gap between error detection and correction.
Meta is in early talks to lease its AI‑compute capacity to Anthropic in a deal that could total up to $10 billion over two years, with monthly payments and an early‑exit clause. The arrangement would let Meta monetize its data‑center assets and diversify beyond advertising, though the talks remain preliminary.
SpaceX is negotiating a deal to lease data‑center capacity worth billions of dollars to the U.S. Defense Department for running AI models. The pact would deepen the contractor's role in Pentagon cloud strategy, pitting it against established players like Amazon and CoreWeave and sparking worries about over‑reliance on a single supplier.
Frontier AI firms publish wildly different capability thresholds, crippling third‑party verification and fostering a race to the bottom. This paper proposes a concrete methodology to derive harmonized minimum standards for misuse risks and AI R&D, rooting thresholds in expected harm and observed progress rates while exposing empirical gaps.
Responsible AI practices have built robust risk mitigation, but the market still fails to reward genuine trustworthiness. Papakonstantis argues that without an independent, outcome‑based certification regime, safety, fairness, and oversight remain invisible to users, regulators, and investors. Introducing such certification would make trustworthy AI measurable, comparable, and commercially valuable.
The scroll‑world skill builds a scroll‑scrubbed, fly‑through 3D landing page for any brand or industry, hookable from Claude Code, Codex or any SKILL.md‑compatible agent. It orchestrates Higgsfield‑powered isometric assets and a vanilla‑JS scrub engine, delivering immersive pages without a custom stack.
Subscribe free