Research notes39 min read
CCAA: AI Rewrites the Standard, Not Just the Job
Make blue dots rare and people start calling purple ones blue. An oystercatcher will abandon its own egg for a more egg-like fake. AI need not replace you; it only has to nudge "normal" upward a few dozen times a day. We name that half CCAA and give an experiment that could start tomorrow.
TopXEA Trading Desk

Contents
- The half nobody measures
- Two kinds of conditioning
- This is not new
- The mirror answers back
- Three mechanisms
- A compiler is not a talisman
- The eight it cannot code
- How solid the flagship case is
- How we would know we were wrong
- Two questions
- Appendix I: On the term
- Appendix II: The coding manual
- Endnotes
- Principal references
In 2018 a group of psychologists sat people in front of a screen full of dots and asked them to pick out the blue ones. Then, quietly, they made blue dots scarce, and the participants began calling purple dots blue. Swap in faces and the same thing happens: make threatening faces rare and neutral faces start to look threatening. Swap in moral judgements and it happens again: make egregious requests rare and innocuous ones start to look unethical. Telling people in advance that this was going to happen made no difference. Paying them to hold the line made no difference either. Levari and his colleagues reported all of that in Science. The point was sixty years old before they made it: Helson's version is that people take their ruler from whatever they have been looking at lately, and never from anything fixed.
Now a bird. Tinbergen gave an oystercatcher a fake egg, larger than its own and more neatly marked, and the bird abandoned its own clutch to sit on the counterfeit. Nothing had gone wrong with the bird. Its instinct was doing exactly what instinct is for, which is to go for whatever is more egg-like. Until somebody built one, nothing in the world had ever been more egg-like than an egg. An exaggerated version of a signal beats the real thing; Tinbergen established that decades ago. The machines this article is about are in the business of building that egg.
Dozens a day, built to each person's taste, and free.
Put a face where the blue dot was, and a conversation where the fake egg was. The retouched selfie in the camera roll is more even than the face in the mirror; the assistant still answering messages at three in the morning is more patient than anyone within reach. Not one of these things is handed over as a finished product. Each one is a ruler, and you hold yourself up against it for years, dozens of times a day, at a cost too small to count. Clinic consultants describe the scene: a patient slides her phone across the desk, and on the screen is an AI-retouched version of her own face, so the question stops being what she wants and becomes whether that face is physically achievable.
Compare for long enough and the more even, better, easier version settles into place as normal, and normal is where everybody starts from when they grade somebody else. We have a name for this class of thing: Cognitive-Conditioning AI Applications, CCAA. It is not the same as AI substitution. Substitution changes who does the work. Conditioning changes what counts as normal work. Swap the ruler and everything moves with it: what clients ask for, what they will pay, how tired the person doing the work feels.
None of this appears on a layoff notice, and none of it enters the unemployment series. You can keep the job and do the work exactly as you did it last year, while the ruler your patients, clients, students and family hold you to was swapped out last night, on a phone. So when you ask whether a job is safe, "can AI do it" is only the first half of the question. Every ruler now used to measure AI's impact measures only that first half. Nobody is measuring the second. The cabinet is full of instruments, and not one of them is pointed at it.

The half nobody measures
The labour market is already registering the first half. On freelance platforms, monthly earnings in highly affected occupations fell by about 5.2% after ChatGPT's release, with no evidence that a strong track record protected anyone.[1] Among workers aged 22 to 25 there is a relative employment shortfall of about 19% in the most AI-exposed occupations against their less-exposed peers, operating through reduced hiring rather than layoffs; the authors themselves call these early descriptive indicators, not causal estimates. In customer support, an AI assistant raised issues resolved per hour by 14%, by 34% for novices, and by almost nothing for experienced high performers. What the machine lifts is the floor.[2] Do not, though, treat ChatGPT as a clean natural experiment. Frank and colleagues, working in 2026 with unemployment-insurance administrative records and millions of LinkedIn profiles, found that unemployment risk in AI-exposed occupations began rising in early 2022, months before ChatGPT was released. The weakness after 2022 is tangled up with monetary policy and post-pandemic sectoral adjustment, neither of which has much to do with what models can do.
The substitution half is real, it is small, and people are already measuring it. The most-cited estimate is Eloundou and colleagues' in Science: counting large models alone, about 1.8% of occupations have more than half their tasks affected; count the software built around them and the figure becomes just over 46%. Change the definition and the answer moves by a factor of twenty-five, and where the difference sits is itself a statement of what these instruments measure, which is whether a task can technically be done. Felten, Raj and Seamans' AIOE ability-linkage measure, and Webb's working paper, still unpublished, measure the same thing.[3] Not one of them contains a variable for who adjudicates whether the work was done well. Not AIOE, not Webb, not Eloundou. They were built from the start to answer a different question. Nobody building them left a column for the person who nods when the work is done, or for where the ruler in that person's head came from.
Mainstream economics has brushed against this axis once. Acemoglu puts AI's ten-year total-factor-productivity effect at no more than 0.71%, and below 0.55% once hard-to-learn tasks are counted in, his stated reason being that such tasks have "many context-dependent factors affecting decision-making and no objective outcome measures from which to learn successful performance". That is very nearly the second axis we want, except that he uses it to argue AI will be less useful and we use it to ask who is protected. Autor puts it more bluntly: the gate around expert work is licensure and liability, and task difficulty is not part of it.[4]
Two kinds of conditioning
For the name to be any use, whatever it does not cover has to be cut away first, or it will swallow everything. The definition is one sentence long and hard going. Every qualifier in it is there to keep something out, starting with AI output that really is handed over as a finished product.
Cognitive-Conditioning AI Applications (CCAA): a class of AI applications whose output is not consumed as a finished product but adopted by the user as a frame of reference. Under long-term, high-frequency exposure at near-zero marginal cost, the statistical properties of that output (better mean, smaller variance, less friction) are internalised as "normal", rewriting the user's expectation threshold for reality.
The trouble is that conditioning has two meanings, and blending them destroys the argument. One of them happens where the work is bought: the buyer's expectations of price, volume, turnaround and revision count are reset. Copywriting is the textbook case, where "a campaign" now means forty variants overnight at zero marginal cost. That is a market phenomenon, standard inflation of the ordinary kind, driven by costs that collapsed and fully explained by price competition, and it needs no new concept. The other happens behind the eyes: a person's baseline for what a face, a conversation, a partner, a body should be is reset. That one is psychological. CCAA refers only to the second.
The exclusion has an uncomfortable consequence. Copywriting, generic illustration, template video — the industries most people reach for first when they think about AI's impact — are precisely where the claim cannot be tested. In those industries the standard moved and the supplier changed at the same time, so nothing can be pinned on conditioning; price competition already explains all of it. Clean evidence exists only where the human is still doing the work and the standard moved anyway: aesthetic medicine, counselling, nursing, teaching. Copywriting cases are used from here on as illustration, never as evidence.
The name is ours, first published on 1 September 2026. "Conditioning" is used here in its everyday sense, not in the sense associative learning has in psychology; and CCAA, as four letters, is a piece of Canadian insolvency law, which makes it close to useless as a search term. In formal writing, spell it out.
This is not new
None of this is new. The blue-dot experiment has a long line of ancestors. Helson said judgement is made against a reference point pooled from whatever has been seen recently. Parducci put numbers on it: the mean and the variance of a sample set the threshold. Brickman and Campbell showed that a better mean becomes the new normal while satisfaction resets to where it started. Behind those stand Festinger's 1954 social comparison, Gerbner's cultivation theory, and the supernormal stimulus that fooled the oystercatcher. Between them they cover the whole of CCAA's mechanistic claim. We do not claim to have discovered a new psychological mechanism.[5]
One more thing has to be said plainly: that experiment ran only half the course. Levari and colleagues ran the scarcity direction, in which blue dots thin out and the concept of blue expands. We need the other one, in which good faces, good replies and good prose grow dense and the threshold for good rises. Reference-point theory is symmetric in principle, but the published experiments went only one way. That symmetry is a hypothesis, not a finding. Testing it is not hard; run the same paradigm backwards.
One item in that same lineage can be picked up and used against us. A central result of the hedonic-adaptation literature is that adaptation is broadly symmetric and reversible. The pleasure of a pay rise decays back to where it started; the treadmill spins, but it does not relocate the runner. On that account, eyes that have grown used to AI-retouched faces will in time grow used to unretouched ones again. If conditioning is only hedonic adaptation under a new name, its economic predictions fail in a specific way: the threshold drifts up, the dissatisfaction fades along with it, and demand does not durably migrate into aesthetic medicine. We do not dispute the objection.
The answer is to narrow the persistence claim. People adapt back; institutions do not. Once a drifted standard is written into a hiring rubric, a clinic's before-and-after gallery, a platform's default photo norm, a performance-review template, an insurance underwriting standard, it is fixed. The gallery is the easiest of these to see. Last year's "after" shots, if they were themselves AI-assisted, are next year's face in the client's head when they walk through the door; nobody will remember which year, or which image, the line started from, and the image is still on the wall, with a date on it. That is exactly what makes such traces measurable, since institutional artefacts are archival, dated and countable. Conditioning produces durable economic consequences only where the drifted standard becomes institutionally encoded.

The mirror answers back
Only one thing here is genuinely new: the marginal cost of generating a better sample has fallen to zero. The mirror now answers back, and it answers one person at a time. The first consequence is that the idealised stimulus is generated per person. Broadcast media produced a single ideal and sent it to everybody, and anyone could see it had been made for somebody else. Generative AI's version is produced only after being optimised against this particular person's responses, and it replies. The social-comparison literature lists contingency among the variables that drive internalisation, and a magazine cover never responded to anyone; it hangs there, the same for everybody. The second consequence is closer to the skin: the reference set now includes the user's own idealised image. Filters and AI self-portraits collapse the comparison distance to zero, and the target is no longer a celebrity but a better version of oneself, with the same bone structure, the same moles, still wearing yesterday's shirt, only with much cleaner skin. One person can be handed dozens of samples optimised for them in a single day, and that number is itself countable.
The casualty is cultivation theory's precondition. Cultivation assumes a shared, broadcast, non-elective symbolic environment, and it is precisely that sharedness which produces mainstreaming, the convergence of everyone's worldview toward the televised middle. Generative AI supplies a personalised, on-demand, unbounded, user-steerable distribution. This is not merely a change of medium; it dismantles cultivation's generative condition. Out of that comes a falsifiable prediction: CCAA predicts normative divergence, not normative mainstreaming. If the coming decade shows convergence in aesthetics, language and relational expectations, this article's central distinction is wrong.
Three mechanisms
Standard drift gets from the screen into somebody's expectations by three routes: the words people write, the range of things they see, and what they come to expect of another person's patience. The three are not equally well evidenced, and the strongest evidence comes first.
Two memos
The first is register assimilation. One request, two memos, two ways of saying it:
- "The following two items require your decision."
- "Mind casting an eye over these two?" / "Sorry to bug you — could you make the call on these?" / "These two need your signature."
Neither is more professional than the other. The difference is that one of them tells you something about the person who wrote it. Register carries relational information: who is asking whom, on what terms, how urgent this really is, how large a favour is being requested. A manager reading the first sentence learns that two things need doing. Reading the second, they also learn how well the writer knows them and how hard they are being pushed. AI writing tools are replacing the second kind with the first, an institutional neutral register.
The second beat comes later. Three requests get polished into the same shape, the manager cannot tell which one is genuinely urgent and ends up calling all three people into a meeting. The words were saved; the work was not. Total communication cost has not fallen. It has moved out of the language and into the process, as another meeting, an earlier deadline, more people on cc.
Start with a causal result from the laboratory. Jakesch and colleagues had some fifteen hundred people write with the help of an opinionated language model. What they wrote was pushed toward the model's position, and the attitudes they subsequently reported moved with it, an effect the authors call latent persuasion. The authors also warn that the model's opinion was deliberately made strong, that the topic was low-entrenchment, and that the attitude shift was likely transient.[6]
Scale the same thing up to an entire scholarly corpus and it is still there. Kobak and colleagues went through more than 15 million PubMed abstracts and, from an abrupt post-2022 spike in excess word frequencies, inferred that at least 13.5% of 2024 abstracts had been LLM-processed, rising to 40% in some subcorpora, a shift larger than the one COVID produced in academic writing. Liang and colleagues went through nearly a million papers and put LLM-modified text at up to 17.5% in computer science.[7]
Now put it between two people. In two randomised experiments, Hohenstein and colleagues found that smart replies raised efficiency and positive emotional language and improved how the conversation partner was rated; but as soon as the partner suspected AI, the ratings turned negative.[8] Laboratory, corpus and conversation all point the same way, and the evidence chain here is the most complete.
Complete is not the same as on point. These studies establish imitation and diffusion: language really is converging on model style. They do not establish that anyone's acceptance threshold has moved. Imitation and a moved threshold are two different things, and the frozen-text experiment further on exists to pull them apart.
How wide normal is
The second is distribution replacement. People judge what counts as normal by estimating a distribution from the samples they have seen. AI-generated content moves two parameters of that distribution at once: the mean rises and the variance is compressed. The compression is the more insidious of the two and the more consequential. It erases the information about how wide normal is, and that is the information which decides how much a person will count as normal at all.
The hardest supporting evidence comes not from AI research but from work on visual adaptation. Brooks and colleagues, reviewing that literature in 2020, summarise experiments showing that prolonged exposure to photographs of bodies of a given size shifts the observer's point of subjective normality, the place where a normal body sits for them. The shift is bidirectional, experimentally manipulable and modulated by attention, and it is the best-evidenced version we have of the claim that a distribution sets the norm. Note the boundary: the stimuli were photographs, not AI output.[9]
On the AI side there is one pre-registered experiment. Doshi and Hauser recruited 293 writers and 600 evaluators; those given generative-AI story ideas produced work rated about 6.7% more novel and 6.4% more useful, with the largest gains going to the least creative writers, while their stories were significantly more similar to one another than the human-only control group's. The magnitude of that similarity result has to be stated plainly: the direction is solid; the size is small.[10]
A year later somebody ran a larger experiment, and the conclusion turned around. Ashkinaze and colleagues reported a dynamic chained design at the ACM Collective Intelligence conference in 2025, with more than 800 participants across more than 40 countries, each wave able to see the previous wave's output. High AI exposure did not change individual creativity, and it raised the diversity of collective ideas and the rate at which that diversity changed. That design is larger and more ecologically valid. This mechanism is contested, not settled. We keep it anyway, on the grounds that both creativity experiments measure output diversity while what we are talking about is variance compression; the evidence from visual adaptation for variance compression is solid, and the creativity work may well not be touching it at all. Keeping it is a bet of ours, and the literature is not currently on our side.[11]
Patience at three in the morning
The third is frictionless reinforcement. This is what an AI assistant is to every one of its users: it answers at three in the morning, it does not sigh at the seventh revision, and it does not hold it against you when you snap at it. Infinitely patient, available at any hour, entirely centred on the user, never talking back, never bearing a grudge, never tired. Historically only a very small number of people ever experienced a relationship of that shape, and societies kept special words ready for it: court physician, private tutor, valet. It is now everybody's daily experience, dozens of times a day. If reference-point theory is right, that daily experience should leave traces: more patience expected of service workers, lower tolerance for friction with colleagues and family, "handled by a human" moving from the default to a paid upgrade, a lower threshold for escalating a complaint. The list is inference, not a verified conclusion.
Conditioning requires repetition, immediate feedback and low cost. Feedback from people always carries friction: you wait for it, you read the face it arrives on, you owe something back for it. AI supplies all three conditions and none of the friction. When the frictionless side becomes the main training set, tolerance for friction should fall. That is the reasoning. The evidence is a good deal softer than the reasoning.
Fang and colleagues had 981 people talk to ChatGPT for four weeks, more than 300,000 messages in all. Nothing came out of the conditions they had arranged, text or voice, personal topics or impersonal. What they did detect was something they had not arranged: the more someone talked to it per day on average, the lonelier they were, the less they socialised offline, the more emotionally dependent they became and the less able they were to stop. Talking more was these people's own choice, so this is correlation, not causation, since usage was self-selected; the effects are very small, with β between 0.02 and 0.06; and what these people used was general-purpose ChatGPT with an assigned daily task, not a companion app. The paper is a preprint and not peer-reviewed.[12]
Then the evidence turns around. De Freitas and colleagues, across several studies in the Journal of Consumer Research in 2026, find that AI companions reduce loneliness, at a magnitude comparable to human interaction and above activities such as watching video. Smith, Bradbury and Karney, reviewing the field from relationship science in Perspectives on Psychological Science in 2025, put it more directly: the empirical base for downstream effects on human relationships is largely absent.[13] Neither side measured the thing that matters. No experiment has yet measured what happens to a person's expectations of human partners or human service workers after exposure to AI. This mechanism is a hypothesis, not a finding, and the weakest-evidenced of the three. We keep it because the other two ask how content changes, and only this one asks how people's expectations of other people change, which is the step that turns standard drift into money.

A compiler is not a talisman
A plumber's work has to pass inspection, and the inspection does not care how nicely the plumber smiles. A metrologist's work has to pass calibration, and calibration does not care what mood the client is in. Where an outsider can enforce a gate, nobody's frame of reference quietly moves the standard. A hard acceptance test means safety: we believed that rule at the outset too, and thought it close to self-evident. The junior programmer overturned it. They have a compiler, they have unit tests, they have production incidents, three gates each harder than the last, every one of them a genuine external acceptance test; when they get it wrong the machine says on the spot which line it was, with nobody's likes or dislikes involved. By the rule above they ought to be safe. They are not. A machine passes all three gates more cheaply than they do. That rule is false. Acceptance did not protect them; acceptance became the entry pass for the machine that replaced them. A gate that scores itself has already written the exam paper for the machine, and handed over the answer key with it.
Turn it around and it comes out right: an acceptance test protects a practitioner in proportion to what it costs the automating system to pass it. Compilers, unit tests, actuarial reserve tests, standardised examinations are cleared by a machine at or below human cost; those anchors are soft and provide no protection. Site inspection, malpractice exposure, structural sign-off require a body present, a licence, somebody with their livelihood on the line; those anchors are hard, and those do protect.
Measuring an occupation therefore comes down to two questions. The first is acceptance anchoring (AA): when the buyer says "not good enough", can the dispute be settled without reference to the buyer's comparison set? For the plumber the answer is yes, since whether the pipe leaks is something an outsider can see. For aesthetic medicine the answer splits in two: the safety half can be settled, and the looks-good half rests entirely on the face in the client's head. The second is deliverable automatability (DA): can a current model produce the accepted deliverable end to end, the deliverable rather than the tasks? Retouching, yes. Cutting, no.
Each question splits into four checkable binary items scored 0–4, with the scoring detail set out in the coding manual at the end. Cross the two axes and four cells appear. This 2×2 is the skeleton of the whole framework, and every judgement that follows is read off it.
| Anchored (AA high) | Elastic (AA low) | |
|---|---|---|
| Automatable (DA high) | Anchored but squeezed (soft anchors fail: junior devs, actuaries, general translation) | Elastic and automatable (content farms, template design) |
| Not automatable (DA low) | Insulated (plumbing, metrology, funeral services, nursing) | Elastic but not automatable (aesthetic medicine, counselling, companionship, luxury experience) |
Look at that diagram and the natural next move is to add a third axis: how far this trade depends on aesthetics, emotion and the body. We drew it that way in our first version, as a second matrix. The professional athlete punched a hole through it single-handed. On dependence on emotion and the body, professional sport is at the ceiling; but its adjudication is the scoreboard, one of the hardest external acceptance tests in existence, and no amount of conditioning in the crowd changes the score. The same object lands at opposite extremes under the two definitions of the vertical axis, which shows that dependence cannot stand in for adjudication mechanism. Nurses and reconstructive surgeons have the same shape. Our own second matrix is wrong on this point.
That diagram did catch one thing this 2×2 cannot see: who manufactures the frame of reference. It has to be drawn, but it cannot be drawn as an axis, because it is a directional flow and an axis carries only magnitude. The wedding photographer shows why. This year's portfolio is next year's bride's reference; the bride arrives asking for a result that is physically impossible and pays for it; and meanwhile their own retouching line has already been automated. One person stands at the emitter, the beneficiary and the squeezed positions at once. Role travels with the revenue line, so it can only be coded per revenue line.
| Loop position | Meaning | Typical revenue lines |
|---|---|---|
| Emitter | Its output circulates as other markets' reference material | image models, beauty-filter platforms, VTubers, AI short drama, companion apps |
| Beneficiary | The drifted standard exceeds what the cheap channel can physically supply, so demand migrates here | aesthetic medicine, high-end spa, counselling, companionship services |
| Squeezed | The standard drifted and the deliverable is automatable | illustrators, junior developers, general translators |
| Insulated | Anchored acceptance; not a participant in the reference-frame game | plumbing, metrology, funeral services, lift maintenance |
So the treatment is to keep one 2×2 and render what the second matrix saw as arrows between the cells. Emitters manufacture the frame of reference; the frame strikes beneficiaries and the squeezed at the same time, demand rising for the first while for the second the standard rises and their own work is replaced. The insulated are not on this chain, and no arrow touches them. The four positions are four links in one supply chain, each link's output the next link's input. Merging the two matrices adds this set of arrows and nothing else.
The eight it cannot code
The framework was stress-tested against 18 industries, and 10 of the 18 rows carry a warning flag. In 8 of those 10, about 44% of the whole table, the coding itself is unstable: aesthetic medicine, tattooing, VTuber agencies, radiologists, wedding photography, actuaries, personal trainers and signature architects. The other two code stably and are troublesome elsewhere: the junior programmer overturned the rule above, and companionship services have a sign nobody can pin down. Where the coding misses is where the evidence stops.
Start with aesthetic medicine, the flagship case and, of all things, the worst-coded of the lot. The safety half scores 4: when something goes wrong there is a licence, a chart, liability, money to be paid out, and precedent for how much. The satisfaction half scores 0: whether it looks good has no external adjudication whatsoever, the only judge being the face in the client's head, and that face was edited on a phone last night. Same clinic, same operation, and the two halves are pulled in opposite directions, the safety half disciplined more tightly every year, the satisfaction half running further out of bounds. The half the surgeon can demonstrate and the half the client cares about lie side by side on the same consent form, neither with any purchase on the other. To give the trade a cell, you have to decide first which half you are listening to.
The next four break in two pairs. VTuber agencies could be replaced by AI outright, with nothing in the law to stop it, and fans pay a premium anyway for the fact that there is a person behind the identity. Signature architects have an elastic standard and a highly automatable deliverable, so the framework says squeezed, and prestige protects them regardless. When one patch fixes two cells, what is missing is a variable. The variable is the provenance premium: somebody is willing to pay for the fact that a human made it. Whether it is stable is an open empirical question the matrix cannot answer.
Radiologists and actuaries come apart somewhere else, at the point where what stops the machine first is permission rather than capability. Model reading of scans and model reserve testing are already at the door; the regulations still want a named person to sign. Cells tagged gated flip fastest and carry the highest variance, so they can only be flagged separately and never averaged in.
That leaves companionship services, which are the framework's hard boundary. Whether AI companionship is a complement (raising the affective baseline and pushing people toward human services) or a substitute (satisfying the need in-channel) is not determined by the two axes; it is exogenous. For high-conditioning, low-automatability services, therefore, we give magnitude but not direction, until complementarity is measured separately.
| Industry | AA | DA | Loop | Stability |
|---|---|---|---|---|
| Aesthetic medicine | safety 4 / satisfaction 0 | 1 | Beneficiary | ⚠ flagship case, worst coding |
| Tattoo artist | 1 | design 3 / application 0 | Emitter+Beneficiary / Squeezed | ⚠ dual role, must split |
| Plumber | 4 (hard) | 0 | Insulated | stable |
| Dentist | restorative 4 / cosmetic 1 | 0 | Insulated (+ beneficiary tail) | stable |
| VTuber agency | 0 | 3, but provenance premium | Emitter | ⚠ see below |
| Junior programmer | 2–3, soft | 4 | Squeezed | ⚠ refutes the old rule |
| Radiologist | 4 (hard) | capability 4 / gated | pending | ⚠ highest variance |
| Translator | certified 3 / general 0 | 4 | Gated / Squeezed | stable |
| Wedding photographer | 0 | capture 1 / edit 4 | all three at once | ⚠ must split |
| Nurse | 4 (hard) | 0 | Insulated | stable; see prediction below |
| Luxury concierge | 0 | presence 0 / information 4 | Beneficiary / Squeezed | stable |
| Actuary | 3–4, soft | gated | pending | ⚠ soft-anchor risk |
| Personal trainer | 1 (arguably 2) | programming 3 / presence 0 | Beneficiary / Squeezed | ⚠ contested-band case |
| Funeral director | 3 | 0 | Insulated | stable |
| Signature architect | structural 4 / design 0 | concept 4 / stamp gated | Emitter+Beneficiary | ⚠ prestige inversion |
| Companionship services | 0 | 0 | Beneficiary | ⚠ sign indeterminate |
| Professional athlete | 4 (hardest) | 0 | Emitter | stable |
| Content-farm writer | 0 | 4 | Squeezed | stable but over-determined |

How solid the flagship case is
For the beneficiary side to hold at all, somebody has to be paying at that end. Aesthetic medicine is the best-documented stretch of the line. Its evidence is weaker than most people assume, and the structure inside it is messier than the usual account suggests.
The strongest single citation available on this subject is a three-level meta-analysis published in Mass Communication and Society in 2026. The correlation between social media use and consideration of cosmetic surgery is r = 0.21, moderated upward by appearance-focused use. It is a correlation, and small-to-moderate in size.[14]
Industry data supplies an inconvenient detail. The ISAPS Global Survey for 2024 reports about +42.5% growth over four years, but 2024 fell against 2023: surgical −6.7%, non-surgical −3.1%. If conditioning-driven demand were monotonic, that decline would need explaining.[15] Our reading is that the turn in that particular year was mainly about price and credit, and that reference-frame drift is a slow variable, too slow to show up in a single year's curve. We cannot prove that reading: on the 2024 data it looks identical to "conditioning was never pushing demand in the first place". Only one thing is certain: "AI rewrites beauty standards, therefore aesthetic-medicine demand explodes" is too simple a narrative, and macro cycles, price and regulation plausibly dominate reference-frame drift.
One attribution needs correcting. "Snapchat dysmorphia" was not coined by JAMA Facial Plastic Surgery. Its first appearance in the peer-reviewed literature is an editorial by Ramphul and Mejias in Cureus in March 2018; press attribution for the coinage goes to the British cosmetic doctor Tijion Esho. The Rajanala, Maymone and Vashi piece is a Viewpoint, and what it did was medicalise an already-circulating term and popularise it; its own wording is "dubbed". This corrects an error in an earlier version of this article.[16]
The 55% figure needs discounting. "55% of surveyed surgeons had seen patients seeking procedures to look better in selfies, up from 42%" comes from an AAFPRS annual member-survey press release; the sampling frame is association member surgeons responding voluntarily, and the sample size, response rate and sampling method have never been published, nor does any corresponding peer-reviewed document exist. Usable as evidence that the phenomenon occurs; entirely unusable as a population rate.
"Emotional labour gets more expensive" currently has no evidence behind it. Hochschild's The Managed Heart (1983) established the concept, and rising customer incivility has peer-reviewed support, with comparisons before and after COVID showing the indirect effect of customer incivility on performance via emotional exhaustion becoming more pronounced. But we could find no study establishing that exposure to technology or to AI raises customers' expectations of human service providers. This inference can only be offered as a hypothesis with a research design attached, never as an established finding.
The same drift shows up on a map, with a lag. Where adoption is dense, the new demand already looks roughly like this: requests to turn a filtered face back into a real one; weaning off AI companionship and repairing human relationships; a premium for human-written content and verified-human identity; expectation-management training for practitioners on the beneficiary side. Where adoption is thin, the insulated cell has not been repriced at all. All of this rests on an untested assumption, namely that CCAA spreads at the same rate as AI adoption. That assumption is quite possibly wrong, because cognitive effects travel through content, and content circulates nationally rather than down a geographic gradient. Use it as a hypothesis, not a plan.

How we would know we were wrong
We wrote three falsification conditions of our own, and thought each one through; all three fail. The first was to count the AI share of the reference images that aesthetic-medicine patients bring in. It has no denominator, no baseline and no threshold to begin with. Clinics do not archive reference images, and consent blocks retrospective collection; even with the images in hand, AI detection on compressed social photos is too poorly calibrated to use. In any case the AI share of the entire image supply is rising regardless, so seeing that patients bring AI-retouched photos would prove only that AI-retouched photos exist.
The second was to ask heavy users of companionship apps whether their tolerance for friction had fallen, and there are walls on three sides. Adopters self-select on the very slope being measured, so controlling for baseline loneliness does not fix it. Ask them to report it themselves and the ruler they report with has already moved along with the thing being measured. Even the sign is undetermined, since therapeutic use might perfectly well raise tolerance.
The third was to measure register change in workplace writing. That can be measured, but it measures something else. Reverse causality, same-window confounds and no control group are all present, and what comes out at the end is imitation and diffusion. Of the three mechanisms, register has the thickest file of material, and it is stuck at this same point.
The design that does work is this. Take 40 human-written emails and memos from 2019 and change not a word. Each year afterwards, a fresh cohort of managers who have never seen them rates those same texts against the same rubric, judging each one "acceptable to send" or "would demand revision". Because the stimulus is constant by construction, any drift can only come from the standard itself.[17]
That design walks around every place the other three were blocked: the stimulus does not move, only the ruler does. It is the line between a framework and a finding. The same design ports, with frozen face panels for aesthetic medicine, frozen conversation transcripts for companionship, frozen text panels for the workplace, and a dose-response test across domains of differing AI exposure. Nor is it expensive: one round a year, forty documents, and a set of raters willing to sit down and read them through. Until those raters hand in their sheets, every quantity in this article is a hypothesis.
There is a cheaper one still, already sitting in the 18-row table, and it is nursing. Nursing behaviour is auditable and roughly constant: how many times a patient is turned, how many seconds until the call bell is answered, how many checks before a drug is given. These things are recorded, and they do not vary by year. If patients' frames of reference are being reset by an infinitely patient chatbot, then with nursing behaviour unchanged, patient-satisfaction scores (HCAHPS-type instruments) should decline year on year. Constant behaviour with a drifting rating is a threshold measurement lying ready to hand. The score tables are public and anyone can pull them; not even forty letters need collecting. All it needs is one external event: patients going and using a chatbot themselves. That is happening, and it is happening at different speeds in different regions and different age groups, which supplies a natural dose gradient. If nursing behaviour has not changed and nurses' scores do not move from year to year, this article is wrong.
Two questions
The framework does not predict employment levels. An industry's employment outcome is dominated by demand, demographics, price and regulation, and the conditioning effect is a small residual; using this 2×2 to forecast how many posts a trade will have in a few years' time is a misuse. It predicts the direction of standard drift only, and durable consequences only where institutional encoding occurs. For high-conditioning, low-automatability services, companionship above all, it can speak to magnitude but not direction.
Two questions are worth carrying away, and you can put both to yourself whatever you do for a living. Who adjudicates whether I did it well, a form with legal force or somebody's feeling? And for a machine to pass that acceptance test, is it dearer or cheaper than me? Answer both and you know which cell you are standing in. Miss the second half and you arrive at "AI can't do my job, so I'm safe", which is what most people on the beneficiary side believe, while the frame of reference they are judged against is being rewritten. Turned around, whether a job is really protected is the same question from the other end: how expensive that acceptance test is for a machine to pass. An acceptance test a machine clears cheaply is not a fence around your trade; it is the syllabus for the thing that replaces you.
The rest is measurement, and it is cheap: one cohort of raters a year, forty pieces of old material that never change, and a score table that is already public. Until the constant-stimulus drift panel returns its first data, this article is a vocabulary and a set of hypotheses. The oystercatcher's trouble was never the bird. It was the egg somebody built for it, and if these machines really are building that egg, the nurses' scores and those forty letters will show it first.
Appendix I: On the term
The term was coined and defined by TopXEA Research and first published in this article, on 1 September 2026. Suggested citation: TopXEA Research (1 September 2026). CCAA: AI Rewrites the Standard, Not Just the Job. TopXEA. https://topxea.com/blog/ccaa-cognitive-conditioning-ai-applications
"Conditioning" is used here in its everyday sense, not its technical one. In psychology, "cognitive conditioning" is already occupied: it belongs to the covert-conditioning family developed by Cautela and denotes associative learning (conditioned stimulus, unconditioned stimulus, reinforcement). The mechanism described here is not associative learning. There is no CS, no US and no reinforcement schedule. It is psychophysical reference-point recalibration: Helson's adaptation-level theory, Parducci's range-frequency theory, Levari et al.'s 2018 experiments. We keep the word because in ordinary language it conveys exactly the right thing, something long-run, slow and unnoticed. In a scholarly context, read the mechanism as described in this paragraph, and not as conditioning in the learning-theory sense.
The acronym collides badly. CCAA carries at least two dozen other senses in other fields, the dominant one being Canada's Companies' Creditors Arrangement Act. That makes CCAA effectively unusable as a search term; in formal writing, spell it out, or write "CCAA (cognitive conditioning)". We claim the definition, not ownership of three English words. If someone has used this phrasing with the same or a similar meaning before us, send us the source and we will add the attribution here and cede priority. A link is enough; no argument is required.
Appendix II: The coding manual
The two axes carry 8 binary items between them, four on each, scored 0–4. The items are written as checkable statements of fact rather than judgements of degree, so that three people coding separately land in the same place. The four items for acceptance anchoring: a written, third-party-enforceable specification exists (statute, building code, clinical guideline, numeric SLA); a pass/fail event is observable by a third party within a bounded time (inspection, assay, whether the pipe leaks); failure carries liability that attaches regardless of satisfaction (licence, malpractice, criminal); payment or renewal is conditioned on that test rather than on a satisfaction rating. Anchored is ≥3, elastic is ≤1, and =2 is the contested band, reported publicly as unstable and never forced to one side.
The four items for deliverable automatability: the deliverable is transmissible as bits, with no physical manipulation of the client's body or property at delivery; ≥70% of the occupation's O*NET task-time items are model-performable at current cost; there is no legal requirement for a human signature or presence, and no measurable willingness-to-pay premium conditioned on human authorship; liability is absent, or insurable without a named human. High is ≥3, low is ≤1, and =2 enters the same contested band. Embodiment goes into item 1 of this axis and nowhere else: it is a reason the deliverable resists automation, and not a reason the standard resists drift. Counting embodiment once is what stops the two axes copying each other's answers. The provenance premium is listed separately as a modifier, and it records whether a measurable willingness to pay attaches to human authorship as such. VTuber agencies and signature architects both draw their revenue from it, and since the property has no place on either axis it hangs outside them, coded and reported on its own. Its value very likely changes over time, so record it with a year.
There are four coding rules. Split revenue lines before coding: if two lines differ by ≥2 on acceptance anchoring or automatability, plot the occupation as a segment rather than a point and report it revenue-weighted, and aesthetic medicine, tattooing and wedding photography all have to be handled this way. When acceptance anchoring = 2, ask who pays for failure, since a defined penalty independent of customer mood codes as anchored while a client who only blames themselves codes as elastic. When automatability = 2, ask what breaks first, capability or permission; if permission, code it low and tag it gated, because those cells flip fastest and carry the highest variance and must be flagged separately rather than averaged in. Re-code the 2×2 once a year and report the delta, the level being a by-product. Automatability is not a constant, and an undated matrix rots.
The reliability protocol runs like this: 8 binary items, 3 independent coders, a 40-occupation sample, a target of Krippendorff's α ≥ 0.7, and publication of the disagreement list rather than a single headline α. Which occupations and which items the disagreements fall on is far more useful than one overall score.
Endnotes
- Hui, Reshef and Zhou, Organization Science 35(6):1977–1989 (2024): on a freelance platform, highly affected occupations saw roughly −2% in work volume and −5.2% in monthly earnings, with no evidence that strong past performance offered protection. Demirci, Hannane and Zhu, Management Science (2025): within eight months, job posts fell about 21% for automation-prone writing and coding work, and about 17% for image work once image-generation AI appeared; the posts that survived were more complex and better paid.
- Brynjolfsson, Chandar and Chen, "Canaries in the Coal Mine?", revised August 2026, using ADP administrative payroll data: a relative employment shortfall of about 19% for workers aged 22–25 in the most AI-exposed occupations. Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 140(2):889–942 (2025): across 5,179 customer-support agents, issues resolved per hour rose 14% on average (note that this is not the widely quoted 15%), 34% for novices, and almost nothing for experienced high performers. Noy and Zhang, Science 381(6654):187–192 (2023), a pre-registered randomised trial: 453 college-educated professionals, writing time down 40%, rated quality up 18%.
- Eloundou et al., Science 384(6702):1306–1308 (2024). A correction in passing to the most common citation error in this field: the widely circulated "80% of workers with at least 10% of tasks affected, 19% of workers with at least half their tasks affected" comes from their 2023 arXiv preprint and not from the Science article; the two must not be cited together. Felten, Raj and Seamans' AIOE ability-linkage measure, Strategic Management Journal (2021). Webb's 2019 working paper remains unpublished.
- Acemoglu, Economic Policy 40(121) (2025); Autor, NBER Working Paper 32140 (2024).
- Levari et al., Science 360(6396):1465–1467 (2018). Sources for the theoretical lineage: Helson, Adaptation-Level Theory (1964); Parducci, Psychological Review 72(6):407–418 (1965); Festinger, Human Relations 7(2):117–140 (1954); Gerbner and Gross, Journal of Communication 26(2):172–199 (1976).
- Jakesch et al., CHI 2023: 1,506 participants; mean difference in what was written 0.29, d ≈ 0.5; attitude shift d = 0.22, p < 0.001 (doi:10.1145/3544548.3581196).
- Kobak et al., Science Advances 11(27) eadt3813 (2025): more than 15 million PubMed abstracts covering 2010 to 2024, inferred by an excess-vocabulary method. Liang et al., Nature Human Behaviour 9:2599–2609 (2025): 950,965 papers.
- Hohenstein et al., Scientific Reports 13:5487 (2023), two randomised experiments, n = 1,036.
- Brooks et al., Perspectives on Psychological Science 15(1):133–149 (2020).
- Doshi and Hauser, Science Advances 10(28) eadn5290 (2024): 293 writers, 600 evaluators; on a 0–100 cosine-similarity scale the homogenisation effect is only b = 0.871 (one AI idea) and b = 0.718 (five); similarity is measured within condition, each story against the centroid of the others in its own condition, not across conditions; the study itself claims no population-level or longitudinal diversity collapse.
- Ashkinaze et al., Proc. ACM Collective Intelligence Conference 198–213 (2025).
- Fang et al. (2025), arXiv:2503.17473: 981 completers, 28 days, more than 300,000 messages, pre-registered on AsPredicted; a 3×3 design of text / neutral voice / engaging voice by open-ended / non-personal / personal topics. Loneliness β = 0.02, p = 0.027; offline socialisation β = −0.05, p = 0.0019; emotional dependence β = 0.06, p < 0.001; problematic use (β = 0.02, p = 0.017).
- De Freitas et al., Journal of Consumer Research 52(6):1126–1148 (2026); Smith, Bradbury and Karney, Perspectives on Psychological Science 20(6):1081–1099 (2025).
- The three-level meta-analysis in Mass Communication and Society (2026): 24 studies, 96 effect sizes, 10,445 participants in total; r = 0.21, 95% CI [0.13, 0.26].
- ISAPS Global Survey 2024: 37.9 million procedures, of which 17.4 million surgical and 20.5 million non-surgical.
- Ramphul and Mejias, editorial in Cureus 10(3):e2263, March 2018 (doi:10.7759/cureus.2263). The Rajanala, Maymone and Vashi piece appeared in JAMA Facial Plastic Surgery 20(6):443–444 (doi:10.1001/jamafacial.2018.0486).
- Statistical specification for the constant-stimulus drift panel: the primary outcome is the year fixed effect on the probability of "would demand revision" across the same unchanged texts, with raters blind and a fresh cohort every year. Identification works off availability rather than usage, using exogenous shocks such as regulator-forced withdrawals and app-store policy changes together with difference-in-differences; an encouragement-design randomised trial, with encouragement as an instrument, also serves. Every self-report scale must be anchoring-vignette-adjusted, or the reference-group effect will contaminate the entire literature.
Principal references
- Levari, D. E., Gilbert, D. T., Wilson, T. D., Sievers, B., Amodio, D. M., & Wheatley, T. (2018). Prevalence-induced concept change in human judgment. Science, 360(6396), 1465–1467. doi:10.1126/science.aap8731
- Helson, H. (1964). Adaptation-Level Theory. Harper & Row. / Parducci, A. (1965). Category judgment: A range-frequency model. Psychological Review, 72(6), 407–418.
- Brooks, K. R., Mond, J., Mitchison, D., Stevenson, R. J., Challinor, K. L., & Stephen, I. D. (2020). Looking at the Figures: Visual Adaptation as a Mechanism for Body-Size and -Shape Misperception. Perspectives on Psychological Science, 15(1), 133–149. doi:10.1177/1745691619869331
- Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. doi:10.1126/sciadv.adn5290
- Ashkinaze, J., Mendelsohn, J., Qiwei, L., Budak, C., & Gilbert, E. (2025). How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas. Proc. ACM Collective Intelligence Conference, 198–213. doi:10.1145/3715928.3737481
- Jakesch, M., Bhat, A., Buschek, D., Zalmanson, L., & Naaman, M. (2023). Co-Writing with Opinionated Language Models Affects Users' Views. CHI '23, Article 111. doi:10.1145/3544548.3581196
- Kobak, D., González-Márquez, R., Horvát, E.-Á., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27), eadt3813. doi:10.1126/sciadv.adt3813
- Liang, W., et al. (2025). Quantifying large language model usage in scientific papers. Nature Human Behaviour, 9, 2599–2609. doi:10.1038/s41562-025-02273-8
- Hohenstein, J., et al. (2023). Artificial intelligence in communication impacts language and social relationships. Scientific Reports, 13, 5487. doi:10.1038/s41598-023-30938-9
- Fang, C. M., et al. (2025). How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study. arXiv:2503.17473 (preprint).
- De Freitas, J., Oğuz-Uğuralp, Z., Uğuralp, A. K., & Puntoni, S. (2026). AI Companions Reduce Loneliness. Journal of Consumer Research, 52(6), 1126–1148. doi:10.1093/jcr/ucaf040
- Smith, M. G., Bradbury, T. N., & Karney, B. R. (2025). Can Generative AI Chatbots Emulate Human Connection? A Relationship Science Perspective. Perspectives on Psychological Science, 20(6), 1081–1099. doi:10.1177/17456916251351306
- Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science, 384(6702), 1306–1308. doi:10.1126/science.adj0998
- Felten, E. W., Raj, M., & Seamans, R. (2021). Occupational, industry, and geographic exposure to artificial intelligence. Strategic Management Journal, 42(12), 2195–2217. doi:10.1002/smj.3286
- Hui, X., Reshef, O., & Zhou, L. (2024). The Short-Term Effects of Generative Artificial Intelligence on Employment. Organization Science, 35(6), 1977–1989. doi:10.1287/orsc.2023.18441
- Demirci, O., Hannane, J., & Zhu, X. (2025). Who Is AI Replacing? The Impact of Generative AI on Online Freelancing Platforms. Management Science. doi:10.1287/mnsc.2024.05420
- Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889–942.
- Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. doi:10.1126/science.adh2586
- Acemoglu, D. (2025). The simple macroeconomics of AI. Economic Policy, 40(121). / Autor, D. H. (2024). Applying AI to Rebuild Middle Class Jobs. NBER Working Paper 32140.
- Gerbner, G., & Gross, L. (1976). Living with Television: The Violence Profile. Journal of Communication, 26(2), 172–199. / Festinger, L. (1954). A theory of social comparison processes. Human Relations, 7(2), 117–140.
- Ramphul, K., & Mejias, S. G. (2018). Is "Snapchat Dysmorphia" a Real Issue? Cureus, 10(3), e2263. doi:10.7759/cureus.2263 / Rajanala, S., Maymone, M. B. C., & Vashi, N. A. (2018). Selfies—Living in the Era of Filtered Photographs. JAMA Facial Plastic Surgery, 20(6), 443–444. doi:10.1001/jamafacial.2018.0486
- Hochschild, A. R. (1983). The Managed Heart: Commercialization of Human Feeling. University of California Press.
Term provenance: the term Cognitive-Conditioning AI Applications (CCAA) was coined and defined by TopXEA Research, first published in this article on 1 September 2026. In this article "conditioning" is used in its everyday sense of habituation and baseline recalibration, not in the technical sense of associative or covert conditioning.
Risk and disclosure: this article is a discussion of a concept and a method. It is not investment advice, career advice or medical advice, and it targets no specific organisation or individual. External research is cited with its source and evidence grade; passages labelled "inference" or "hypothesis" are our judgement and are not verified. Trading carries risk, every EA and quantitative strategy can lose money, and forex and precious-metals trading is high risk — only trade with money you can afford to lose.
Keep reading
Six-fold disagreement on one question — definitional, not a research-quality problem. The alternative: build it bottom-up from three quarterly-filed AI revenue lines.
We parsed all 90 of Vertiv's (NYSE: VRT) Forms 4 over twelve months, 150 transactions, one at a time. Strip out cashless exercises and net accumulators and the usable figure is $100.8m — of which 91.9% came from the board, while the CEO and CFO sold nothing. Reproducible steps included.
For ten consecutive quarters Vertiv disclosed order growth, book-to-bill and backlog every quarter. In Q4 2025 those read +252%, 2.9x and $15.0bn — and on the same call management announced they would no longer be reported quarterly; both releases since count zero. Includes the reproducible method, and the trap that produces the opposite conclusion if you skip it.