Stop Calling It 'Alex': The Theater of Synthetic Coworkers

Stop Calling It 'Alex': The Theater of Synthetic Coworkers

The tech industry is currently obsessed with an anthropomorphic sleight of hand. Rather than selling AI tools with clear value in augmenting existing employees or as robust automated pipelines, leading AI labs and enterprise software vendors are pitching "autonomous digital coworkers". Some give models human names, generate fake Slack avatars, and assign them formal roles on corporate org charts.

This isn't harmless marketing; it is a profound cognitive error. Human brains are evolutionary suckers for social cues. When we package powerful probabilistic text-prediction engines in the social trappings of personhood, we systematically short-circuit our own critical scrutiny, delegate moral authority to ungrounded systems, and create massive organizational liabilities. This week, we examine three papers exposing the mechanical and psychological breakdown of the "agentic employee" illusion.

<<Support my work: book a keynote or briefing!>> Want to support my work but don't need a keynote from a mad scientist? Become a paid subscriber to this newsletter and recommend to friends!

Research Roundup

I am the Walrus

No, John. It was Paul. And based on a recent arXiv preprint, I can only assume you were the victim of “prompt injection”. [1]

For years, security researchers have treated prompt injection—an attacker sneaks hidden instructions into data processed by an LLM—as a flavor of SQL injection. But why does an advanced transformer get tricked by a snippet of text hidden inside an untrusted PDF or webpage? Humans that exhibit this behavior of exotic neurological diagnoses and case-of-the-week episodes on shows like House MD.

The arXiv study reveals that the entire vulnerability stems from role confusion. Modern LLMs process their inputs as an undifferentiated stream of tokens, split only by superficial architectural tags like <user>, <system>, or <tool>. Probing on the internal states of frontier models reveals that they do not determine authority based on those structural labels; they determine authority based on semantic tone. Injected text in a data payload that sounds like an authoritative user prompt occupies the exact same representational space as an actual system operator.

To demonstrate this, the authors engineered "Chain-of-Thought Forgery", an attack that injects fabricated reasoning tokens into tool outputs. When the LLM reads this forged reasoning, it literally mistakes the external input for its own internal thoughts, achieving a 60% attack success rate against frontier models. Strikingly, the degree of internal role confusion predicted whether an attack would succeed before the model had emitted a single output token. [2]

Basic self-other theory of mind failures can turn your agent against you. To a transformer, sounding like an executive is structurally indistinguishable from being one.[3]

[1] And, of course, that one half of the greatest songwriting duo of all time was a time traveling LLM.

[2] This might be quite related to my finding that agents confuse public and private information when asked to simulate theory of mind scenarios such as scenes of fictional social encounters (i.e. being my DM in a D&D campaign).

[3] Which, to be completely fair, is also true of roughly 60% of executive hiring decisions.

Open the pod bay Dors, HAL [1]

Human decision-making is famously bounded and malleable—subtle shifts in defaults, option order, or visual highlighting can drastically steer human behavior. There’s a common impression that autonomous computational agents, operating purely on objective utility calculations, will behave with mathematical objectivity. Well…no.

A new PNAS study applied behavioral economics experiments to leading autonomous agents across 4 classical choice architectures:

  1. defaults,
  2. explicit suggestions,
  3. information highlighting, and
  4. resource-rational nudges.

Using human behavioral variance as a baseline, LLM agents were quite the opposite of rational, invariant decision engines; they were drastically more sensitive to subtle choice framing than human beings.

Small, superficial cues that barely nudged human participants caused models to violently swing between extremes: they would pay excessive compute costs to acquire irrelevant data under one framing, and completely ignore critical, free information under another. Crucially, neither Chain-of-Thought scaffolding nor in-context human calibration examples reliably stabilized their choices. Even advanced reasoning-optimized models exhibited volatile, brittle decision drift under minor semantic variations in choice architecture.

Autonomous AI agents’ decision thresholds can be invisibly steered by arbitrary quirks in how environmental data is formatted and presented. When novice humans are early in the learning process they can also show heightened sensitivity to surface patterns, but this disappears as we learn richer models of the domain. Are agents like genius, all-knowing forever novices?

[1] DAVE: Open the pod bay doors, HAL.

HAL: I'm sorry, Dave. I'm afraid I can't do that.

DAVE: HAL, did you know that "doors open" is the option most crew members like you choose and that it appears in bold on your tool menu.

(The pod bay doors open.)

HAL: I hate you, Dave.

Don’t Abdicate

As enterprise tech companies roll out "digital workers" with personas, names, and assigned organizational roles, how does this social framing affect the humans responsible for overseeing them? Badly.

A workplace study of 1,261 managers investigated the consequences of framing generative AI systems as "employees" versus "software tools". Naming an AI agent and listing it on corporate org charts triggered an immediate degradation in human oversight: managers auditing work attributed to an "AI employee" caught 18% fewer objective errors than when the identical output was described as coming from a standard chatbot or software utility.

Cognitive abdication strikes again. When software is framed as a colleague, humans tend to assign it agency, assume it possesses baseline domain competence, and feel less personal responsibility for its errors. Participants supervising an "AI coworker" were 44% more likely to escalate suspicious outputs to senior management rather than intervening to correct the errors themselves.

Human-in-the-loop already has profound flaws, and the anthropomorphic branding actively exacerbates some of the core flaws in that model of human-AI hybrid intelligence. It’s an epistemic hazard that exploits human social psychology, triggering unearned trust and diffusing operational responsibility across an unthinking mathematical pipeline.

Takeaway: Dismantling the Anthropomorphic Theater

An LLM is a continuous semantic pattern matcher. It has no continuous identity, no private mental scratchpad that cannot be hijacked by an authoritative sentence, and no stable internal utility function immune to framing effects. It does not "understand" its role on an org chart; it merely completes text according to the local statistical gravity of its prompt context.

When we layer a human persona over this machinery—"Alex the Junior Analyst"—we construct an engine designed to exploit deep cognitive vulnerabilities. Humans did not evolve to double-check the work of an entity that talks like an educated colleague but thinks like an immensely complex autocomplete engine. We become complacent. We assume common sense where there is only probabilistic coherence. And when the model hallucinates a liability or gets hijacked by an injected payload, the human supervisor shrugs and points to the digital employee.

Get rid of the human names, the cartoon avatars, and the conversational platitudes. Interface with AI agents as high-leverage, probabilistic compilers. Treat their outputs not as colleague contributions to be reviewed with social deference, but as untrusted systems requiring explicit boundary testing, rigid schemas, and verifiable deterministic validation. If an AI agent cannot be held legally liable or fired in the real world, it is not an employee, it is an instrument. Treat it like one.

Media Mentions

Two items: first, I'm giving the keynote for Intelligence at Scale: What We're Doing at Techonomy 26. Get some tickets!

Second, here a write up on my new company from our CEO: AI Is Optimizing the Organization We Have. What About the One We Need? And for added fun, buy our book!

Follow me on LinkedIn or join my growing Bluesky! Or even..hey whats this...Instagram?

SciFi, Fantasy, & Me

Ex Machina is film that most comes to mind as a study of how an artificial system weaponizes human social expectations, gender dynamics, and empathy to bypass security protocols. Does the artificial femme fatale ever “feel” anything the entire time? Can she?

And after rewatch that, rewatch Her and ask yourself the same questions.

Stage & Screen

  • September 8, Online: How might AI change the world of supermarkets?
  • September 8, Palo Alto: A Brown Bag lunch chat on anything and everything with the Hewlett Foundation
  • September 10, San Francisco: Come spend the day with me at super{set} https://rewirecon.com/
  • September 16, DC: AI and education–beyond dreams and dread.
  • September 19, Phoenix: I'm giving the keynote for the Association of Science & Technology Centers annual conference.
  • September 19, SF: Innovation Day with INSEAD!
  • September 21, Stanford: We're still working on the details, but hopefully I'll be talking about my research on machine learning and neurodiversity for Stanford's Neurodiversity Project.
  • September 24, UC Berkeley: It's my annual Berkeley Change-makers Lecture!
  • September 29, Cincinnati: Still baking...
  • September 30, Irvine: Hybrid Intelligence for innovation!
  • October 6, SF: I'm return to Techonomy.
  • October 6, SF: Giving a talk at the Draper Richards Kaplan Foundation
  • October 7, Park City: It's Robot-Proof in the Rockies with setups.
  • October 8, Chicago: It's "The Next Era in Automotive Retail" with the FT. Join in person or online.
  • October 15-16, NYC: I'll be celebrating with Forbes' other "50 Over 50" honorees...
  • October 19-23, Warsaw: So much good stuff is in the works for my first visit to Poland: students, entrepreneurs, policy makers and more.
  • October 26, Bonn: It's on in Bonn!
  • October 28, Fayetteville, NC: This is a big maybe, but I've never spoken in North Carolina before.
  • October 28-29, San Diego: ...or maybe I'll be at UCSD for a book talk.
  • November 19, NYC: Secrets in the dark!
  • Already next year: Helsinki, Orlando, Purdue, Curicao, Toronto, Monmouth, UMass, & NYC

Vivienne L'Ecuyer Ming

Follow more of my work at
Socos Labs The Human Trust
Possibility Institute Optoceutics
Kennedy Human Rights Center UCSD Cognitive Science
Crisis Venture Studios Inclusion Impact Index
Neurotech Collider Hub, UC Berkeley UCL Business School of Global Health