Ask who actually gains the most from AI and the empirical literature gives you two answers that contradict each other.
One camp finds that novices gain the most. The most cited result is Brynjolfsson et al.'s study of 5,179 customer support agents: after a generative assistant was rolled out, overall productivity (issues resolved per hour) rose 14%, the least experienced cohort improved 34%, and senior agents barely moved. Writing tasks show a similar convergence. The bottom of the distribution gets pulled up and the spread within the group shrinks. On this evidence, AI is an equalizer.
The other camp observes the opposite. In work that requires professional judgment, experts get noticeably more out of the same tool, and the gap widens. One tool, two opposite conclusions, and both sets of numbers are real. The usual move is to pick whichever result matches the story you already wanted to tell, which is not much of a resolution.
What the paper says
A position paper presented at the 5th International Conference on Hybrid Human-Artificial Intelligence (HHAI 2026, Brussels) proposes a reconciliation: both camps are right, and the dividing line is task complexity.
Drawing on structured field observation in enterprise deployments and teaching settings since mid-2024, the author draws a boundary. On routine, well-structured tasks with a clear right answer, AI acts as an equalizer and the novice-expert gap flattens. On complex tasks that demand deep judgment, AI acts as an amplifier.
The mechanism behind the amplification is more interesting than the conclusion itself. On complex tasks the model's output inevitably contains errors, and experts and novices handle the same flawed output in completely different ways. An expert can verify it, push back on it, spot which specific claim is wrong and demand a redo. For that person the model becomes a lever on their own judgment. A novice has no way to verify, so the output gets accepted wholesale and the errors travel straight into the final work product. What AI amplifies, in other words, is not "skill at using AI" but the judgment the user already had, including the absence of it.
Read that way, the two camps stop contradicting each other. They are measuring the same tool against tasks that sit on opposite sides of the boundary, and each is reporting accurately about the side it sampled.
The implication follows directly: what determines who benefits is domain expertise, not prompt engineering. Being good at writing prompts is not a durable advantage. Domain judgment is.
How strong is the evidence
This is a position paper. Its basis is structured field observation, not a controlled experiment. The two camps it reconciles are themselves randomized controlled trials, which means that on the evidence ladder it sits below the work it is trying to explain. Readers should keep that straight: what it offers is an explanatory framework, not a test of that framework.
"Task complexity" is also never operationalized into anything measurable. What counts as routine and what counts as complex is drawn by example rather than by definition, and applied to a specific job the boundary gets blurry fast. Most real roles are a mixture: a few genuinely routine steps, a few that hinge entirely on judgment, and a long middle where reasonable people would disagree about the label. Anyone who wants to use this as a basis for decisions has to supply that step themselves.
For a position paper none of this is disqualifying. It does mean the framework's current status is "a hypothesis worth testing" rather than "a settled result."
What it means in practice
The framework has a direct effect on how we approach enterprise AI delivery. Hand the routine stages to AI to flatten out efficiency; keep the complex judgment stages with experts and use AI to amplify them. The two have to be designed separately. One process cannot cover both, because the thing each stage needs from a human is different: throughput on one side, verification on the other.
It also explains something we run into repeatedly. The same toolset gets rolled out company-wide, some teams get real results from it, and other teams produce worse work than before. The difference is usually not whether the training was adequate. It is which side of the line that role's core tasks fall on. FIM's internal training programs are organized by role and domain rather than as general AI literacy, and this is the reason.
There is a corollary for individuals too. If the durable advantage comes from domain judgment, then time spent chasing tools and prompt tricks has a return that decays quickly, while time spent on judgment in your own field compounds.
The paper appears in IOS Press's Frontiers in Artificial Intelligence and Applications, volume 423, pp. 212–220, open access under CC BY-NC.