Articles
Articles
Equalizer or Amplifier? The Dividing Line Is Task Complexity
A read of an HHAI 2026 position paper: AI narrows the novice-expert gap on routine work and widens it on complex work. What decides who benefits is domain expertise, not prompt engineering.
Read More
The Truth About AI Right Now: An Expert Interview Without the Hype
AI is everywhere right now — on roadmaps, in boardrooms, and across product conversations. But once you look past the headlines, the questions that matter are still very practical: What’s actually working in real deployments?
Read MoreBeyond RAG: Giving a Language Model Working Memory It Actually Manages
A read of the arXiv preprint on Cognitive Workspace, which replaces retrieve-and-forget with active memory management and layered cognitive buffers: 58.6% memory reuse and a 17–18% net efficiency gain. Includes a self-correction written a year later.
Read MoreWhen Should an Always-On Assistant Speak Up?
Survey walkthrough (under review at TMLR). Research on interruptibility from 1999 formalized the cost of interrupting all the way down to utility theory. Proactive agents built since 2024 dropped that term entirely. The survey reconnects the two lines under a single decision rule.
Read MoreWhat Actually Lives in a Model's Workspace: Registers, Not Plans
Paper notes (under review at BlackboxNLP): an independent replication of the "global workspace" claim on open-weight models. Once you filter using the model's own next-token distribution, the only thing left that survives causal manipulation is context registers. Content plans fall to the permutation baseline.
Read MoreAsk the Same Question Ten Times: Consensus Density Decides How Much the Answers Spread
Paper walkthrough (under review at TMLR): the output dispersion of an aligned model is governed by the consensus density of the task. A judge that sees only the prompt, never the outputs, predicts dispersion at rho = -0.91, and alignment turns out to amplify a gradient the base model already had.
Read MoreLong-Conversation Memory: Raw Transcript Chunks Beat Structured Extraction
Paper notes (under review at ARR): a controlled ablation inside a fixed pipeline finds that storing raw conversation chunks outscores LLM-extracted structured memory by 15.9 points on LoCoMo and 22.0 points on LongMemEval-S, at a lower cost per correct answer.
Read MoreContent Poisoning in Chinese AI Search: Detecting It Is Not the Same as Knowing Who Did It
Paper walkthrough (under review): detection and attribution of GEO poisoning against DeepSeek, Doubao and Kimi. Content features catch the manipulation (F1 0.93) but cannot name the actor; account-interaction graphs do the opposite (0.96). The two need separate models.
Read MoreStay Updated
Get the latest technical articles, industry insights, and product updates.
We respect your privacy. No spam, ever.