Paper Walkthrough

Are Machine-Written Papers More Novel? The Gap Sits in One Facet

A read of a NeurIPS 2026 AI4MetaScience workshop poster (non-archival): machine-generated papers are judged facet-novel more often than matched human papers, but the gap comes mainly from the purpose facet, and LLM re-auditors return refutation rates from 0% to 100% on the same verdicts.

Are Machine-Written Papers More Novel? The Gap Sits in One Facet
5 min read

Ask a language model whether a paper is novel and you will get an answer. Ask a second model to check the first and you will get another answer. Whether either answer measures anything is a separate question, and it has become a practical one: the FARS pipeline produced 166 complete papers at a reported cost of $186,000, and the Agents4Science 2025 venue received more than 300 AI-authored submissions.

This paper, a poster at the NeurIPS 2026 Workshop on AI for Meta-Science (non-archival; the citable version is the Zenodo preprint), audits machine-generated papers for novelty one contribution claim at a time. It finds a gap in favor of the machines, and spends most of its length on why that gap should not be read as a conclusion. There is also a 7-minute video summary.

Watch on YouTube

What the paper does

The unit of analysis is a contribution claim. Each paper is split into claims, and each claim into four facets: purpose, mechanism, evaluation, and domain. The machine arm is the 166 FARS papers (549 claims). The human arm is 166 ICLR 2025 submissions matched one-to-one by title and abstract similarity from the 8,613 with a decision (494 claims).

Prior art is retrieved separately for each facet and cut off at each paper's own date, with the same retrieval stack for both arms. Two judges, GPT-4o and Claude Sonnet 4.6, answer one factual question per facet: does any retrieved work already contain this? Gemini 2.5 Pro breaks ties. No model assigns the final label. A rule derives it: any uncovered facet makes a claim facet-novel; all facets covered by one prior work makes it covered; all facets covered only by several works together makes it a recombination. The rubric was calibrated on 20 validation items and frozen before the full run.

Under that frozen rule, 56.1% of machine claims and 35.6% of human claims are facet-novel, a gap of 20.5 points (95% CI 14.3 to 26.5). Accepted and rejected ICLR papers look nearly identical on this measure (34.4% and 36.6%).

Where the gap comes from

In 203 of the 308 machine facet-novel verdicts (66%), and 102 of the 176 human ones (58%), purpose is the only facet the judges found uncovered. Tightening the rule changes the picture:

  • an uncovered facet other than purpose: gap 4.1 points (95% CI −0.3 to 8.5)
  • an uncovered mechanism: gap 1.8 points (−2.2 to 5.9)
  • at least two uncovered facets: gap 5.8 points (2.3 to 9.4)

Keeping only the better-matched half of the pairs leaves a gap of 12.9 points (4.1 to 21.8), and differences in claim type (machine claims lean toward findings, human claims toward methods) account for 2 to 5 of the 20.5 points. What separates the two arms is mostly how their stated objectives compare with retrieved work. On mechanism, they barely differ.

How strong is the evidence

This is a non-archival workshop poster, and no human calibration has been run. The paper itself classifies every novelty rate as an exploratory upper bound and the machine-versus-human gap as a hypothesis. Two measurements support that classification.

The judge layer. An adversarial auditor was given the same retrieved neighbors and asked to refute sampled purpose-novel verdicts: three auditor models, two prompts, the same 20 items per arm in every cell. With the prompt anchored to the panel's coverage standard, GPT-4o refuted 3/20 machine and 5/20 human verdicts, Sonnet 1/20 and 0/20, and Gemini 19/20 and 19/20, and the GPT-4o and Sonnet flags did not overlap. With an unanchored hostile-reviewer prompt, GPT-4o and Gemini refuted all 20 in both arms. On identical items the refutation rate ranges from 0% to 100% depending on model and prompt. A 2-of-3 majority vote follows the prompt (4/20 and 5/20 anchored, 20/20 unanchored), not the data.

The retrieval layer. A gold set of 106 prior-art pairs, built independently of the retrieval channel, shows the deployed retrieval placed the known prior work in the judged claim's neighbor list for only 24.6% of machine and 28.6% of human pairs. Among gold-sample claims judged facet-novel, the gold paper was never retrieved in 7 of 12 machine and 5 of 10 human cases. Machine-arm recall is lower, so part of the gap may itself be a retrieval effect.

A separate integrity audit of 306 Agents4Science 2025 submissions, double-coded blind, found hard evidence of data fabrication in 0 of 47 accepted versus 16 of 197 rejected submissions (one-sided Fisher p = 0.029; 0.072 when only cases flagged by both coders count). The paper reports this as an association resting on a fragile zero cell. In 10 of the 16 rejected cases, the venue's AI reviewers named the fabrication, often by quoting the authors' own answers on the mandatory AI-involvement checklist.

What it means in practice

Zico reviews contracts in two tiers: a fast model scores each clause, high-confidence results are marked directly, and low-confidence ones go to a large-model deep review or to a person. The paper sets two limits on this kind of model-checks-model design.

First, a reviewing model does not calibrate itself. Every review step has a strictness setting fixed jointly by the model and the prompt, and choosing a different pair can move the rejection rate from 0% to 100% on the same inputs. Voting across models removes their disagreement and keeps the bias they share. Pass-through thresholds and deep-review strictness have to be set against human-labeled clauses, and each judgment should be logged with the model and prompt that produced it. A deep review that agrees with the fast tier does not show that the fast tier is right.

Second, "not found" is not "does not exist." In contract review, "no risk found" and "no matching clause found" are negative results too, and their miss rate should be measured on an independently labeled sample before they are presented as findings. At the recall measured here, a novelty verdict means the retrieval did not find the prior work, which is a weaker statement than novelty.

The integrity result points to a process design as well. The mandatory disclosure checklist turned several fabrications into self-admissions that automated reviewers could then cite. Where AI-generated content enters a contract or document workflow, asking for disclosure first and running automated checks second catches more than detection alone.

Further reading

The paper this article covers is listed on the Research page.