Research

Turning frontline deployments into reproducible research.

Human-AI collaboration, LLM memory, interpretability, and AI-search safety: formalizing production practice into reproducible work. Below are the team's public papers, patents, and academic service, each linked to a verifiable source.

Papers & patents

  • HHAI 2026 · IOS PressHuman–AI2026

    AI as Equalizer or Amplifier? Task Complexity as the Moderating Factor for Human Expertise in Hybrid Intelligence Systems

    Tao An

    Published in the proceedings of the 5th International Conference on Hybrid Human-Artificial Intelligence (HHAI 2026, Brussels), IOS Press Frontiers in AI and Applications vol. 423, pp. 212–220 (open access, CC BY-NC). Drawing on structured field observations since mid-2024, this position paper reconciles the 'equalizer' and 'amplifier' debates: AI narrows novice–expert gaps on routine tasks but amplifies them on complex tasks requiring deep judgment. Domain expertise, not prompt engineering, determines who benefits most.
  • TMLR 2026 · SurveyHuman–AI2026

    When Should the Agent Speak? A Survey of Intervention Timing for Always-On AI Assistants

    Tao An

    A survey of intervention timing for always-on assistants (smart glasses, MR headsets, ambient copilots), organized around one decision rule: intervene iff the expected benefit of acting exceeds the expected cost of interrupting. It reconnects two lineages that do not cite each other: the 1999–2017 interruptibility literature, which formalized interruption cost rigorously but had no capable actor, and the 2024–2026 proactive-agent wave, which has actors but rediscovers the cost term only in fragments (false-alarm pricing, cognitive load, social violation, compute) that no work unifies. Evaluation is the gating layer: mobile health already learns intervention timing online against a validated proximal outcome, while open-world assistance has no analogous outcome, so a reinforcement-learning objective for proactivity has no shared, validated reward to train against. Proposes the design of a benchmark for open-world intervention timing with an explicit cost term.
  • NeurIPS 2026 Workshop · AI4MetaScience · Poster (non-archival)Metascience2026

    Recombination or Discovery? A Retrieval-Grounded Novelty Audit of Machine-Generated Research Papers

    Tao An

    Poster at the NeurIPS 2026 Workshop on AI for Meta-Science (AI4MetaScience; non-archival). A retrieval-grounded, protocol-frozen novelty audit of 166 machine-generated papers (FARS) against 166 topic-matched ICLR 2025 submissions. Each paper is decomposed into contribution claims on four facets (purpose, mechanism, evaluation, domain; 549 machine and 494 human contributions); prior art is retrieved under per-paper submission-date cutoffs; contributions are classified as covered, recombination, or facet-novel under a pre-registered two-judge protocol. Machine contributions are judged facet-novel more often than human ones (56.1% vs. 35.6%), and the gap rests mainly on the purpose facet: requiring an uncovered facet other than purpose shrinks it from 20.5 to 4.1 points (95% CI −0.3 to 8.5). The judge layer cannot be certified by another LLM: adversarial re-auditors return refutation rates from 0% to 100% on the same items depending on auditor model and prompt. A 106-pair gold prior-art audit finds the deployed retrieval surfaces the known prior art for only 25–29% of pairs per arm, so automated novelty rates are exploratory upper bounds. A companion integrity audit of 306 Agents4Science 2025 submissions finds hard fabrication evidence in 0/47 accepted versus 16/197 rejected (one-sided Fisher p = 0.029; 0.072 under conservative coding), an association that reflects both AI-reviewer detection and authors' own disclosure.
  • Preprint · arXivLLM Memory2025

    Cognitive Workspace: Active Memory Management for LLMs

    Tao An

    Proposes Cognitive Workspace, which adds active memory management on top of on-demand retrieval, with hierarchical cognitive buffers and task-driven context optimization modeled on human cognition. Reports a 58.6% memory-reuse rate (vs. 0% for RAG) and a 17–18% net efficiency gain.
  • Preprint · arXivLLM Memory2026

    Fidelity Before Structure: Verbatim Chunks Beat Lossy Artifact Extraction in Long-Conversation LLM Memory

    Tao An

    A controlled ablation isolating the stored memory representation inside one fixed retrieve–rerank–reason pipeline: LLM-extracted typed artifacts versus verbatim conversation chunks. Verbatim chunks win by 15.9 points on LoCoMo (43.9% vs. 28.0%) and 22.0 points on LongMemEval-S. Structured memory should augment verbatim text, not replace it. (Formerly titled It's Fidelity, Not Structure.)
  • Working paperHuman–AI2026

    Consensus Density Predicts Output Dispersion in Aligned LLMs

    Tao An, Shuai Feng

    Sampling an aligned LLM repeatedly and embedding the completions, output dispersion tracks the consensus density of the task: near-zero on factual prompts, wide on open-ended ones (Spearman ρ = 0.85), replicating on a second model and predicted by held-out judges that score only the prompt (ρ = −0.91). A base-vs-instruct comparison shows alignment amplifies a gradient the pretrained base already carries.
  • Working paperInterpretability2026

    Registers, Not Plans: What Lives in a Language Model's Workspace That Isn't on Its Tongue

    Tao An

    An independent replication of Anthropic's global-workspace claim on open-weight models. Filtering lens readouts by the model's own next-token distribution splits the workspace in two: causally steerable context registers (the conversation's language, a corrected typo) survive, while content plans (rhyme, arithmetic) fall to the permutation floor. Editing a register also rewrites the model's representation of the question it was asked, stably across a 1.7B–14B ladder and a second architecture.
  • Working paperAI Safety2026

    CGEP: Toward Detecting and Attributing GEO Poisoning in Chinese AI Search

    Tao An

    Defines GEO-poisoning detection and attribution for Chinese generative search (DeepSeek, Doubao, Kimi): a five-technique taxonomy of coordinated inauthentic manipulation, a task reframing from attack-success to detection → classification → account-cluster attribution, and a legally-constructed synthetic benchmark. A provenance pilot shows detection and attribution need different substrates: content features detect that manipulation happened (F1 0.93) but only an account-interaction graph attributes it to a seller cluster (0.96), and a confidence-gated fusion covers the taxonomy where a learned GNN and a zero-shot LLM both fail.
  • Patent · Under examinationKnowledge Graphs2026

    A Graph-Neural-Network Method for Data-Information Recommendation

    Tao An

    Chinese invention patent, under examination. GNN-based recommendation over heterogeneous data–information graphs.

Academic service

  • Ethics Reviewer

    NeurIPS 2026· 2026

    Tao An

    Ethics Review Committee, Conference on Neural Information Processing Systems (NeurIPS 2026), reviewing flagged submissions against the NeurIPS Code of Ethics: data provenance and informed consent, dual-use and misuse risk, human-subjects considerations, and broader societal impact.

Want to go deep on a research direction?

Research collaboration, a technical approach, or just a chat about a paper? Reach out directly.

Contact us