op-ed · research

How we look for knowledge

Someone asked me where conversation design fits next to technical documentation. It sits before the answer: somebody decides in advance what a system may say and where each reply comes from. Checking that against the research took five studies, and one of them corrected a claim I had been repeating for months.

Every number in this piece has a sample size next to it. The first draft had none, and I could not tell which half of it I believed.

The question I was asked

What exactly is conversation design, and where does it fit as opposed to technical documentation?asked in a conversation, july 2026

Technical documentation explains how a thing works, and how to repair it when it breaks. Conversation design decides what a system says back when someone asks it something about your product, and whether the reply can show where it came from.

People ask a product's chat everything. Why should I buy this. How do I troubleshoot the import. How does this technology actually work. This is terrible. How do I make pancakes. The range is the whole difficulty. You do not get to choose what arrives. You choose what goes back, and that means keeping it on the subject, in the company's voice, and technically right.

You calibrate what the system knows about the product first, and that is the documentation. Nothing built on top of it comes out better than what it was grounded on. Then the guardrails: which subjects it takes, what it does when the match is weak, what it declines. Then the design, the turns, the repair. Then the spoken layer, if it speaks, down to the pauses. That order is the definition, and it is the part usually left vague.

So the two jobs share a spine. A conversation designer with no documentation has nothing to calibrate against, and a technical writer whose material never reaches the system has written for nobody.

A machine reads it first

State of Docs 2026 asked people who write documentation how theirs gets found. Search engines were named by 45%, AI search by 35% (ChatGPT, Perplexity, Google AI Overviews), coding assistants by 18% and MCP servers by 16%. At the largest companies, discovery through AI reached 46%.

The survey is weaker than it looks. The answers are multi-select, so they overlap and add up past a hundred. It went to the people who write the documentation, so it measures what writers believe, and it watched no readers. The sample is 1,131 people, and the report publishes no recruitment method and no response rate.

So I take it as direction and nothing narrower. The direction is that the sentence a writer publishes is often not the sentence a reader sees. Something reads the documentation first, and the person gets whatever that something made of it.

We judged before we read

Peter Pirolli and Stuart Card named the old behaviour information scent in 1999, borrowing from research on how animals forage. You decide whether a patch of information is worth the walk before you reach it, using whatever is visible from where you stand: a title, a date, a domain you recognise. Ten results, one glance, decision made.

Every one of those cues is something a writer made. The title, the heading you skim to, the date, the breadcrumb telling you where in the product you are. That is documentation work, and much of it exists to be judged before it is read.

The cues stop arriving

Hannah Kim, Sergei Kosakovsky Pond and Stephen MacNeil went looking for that scent inside generated answers in 2025, and found the usual carriers weak. Bullet points and related terms, normally strong signals, did less work in a generated reply. What told them whether an answer was any good was mostly whether they already knew the subject.

So the cues still get written, and they stop arriving. The title, the date and the structure stay in the documentation, and the reader gets a paragraph with none of them. Judging moved off the page and into the reader's head, which is the one place nobody can design.

The paper is autoethnographic: three researchers watching themselves work in bioinformatics, no outside participants and no sample. I use it as one careful observation and never as evidence about users.

Biggest help, least experience

Erik Brynjolfsson, Danielle Li and Lindsey Raymond followed a conversational assistant as it rolled out across a customer support organisation, 5,179 agents in all, and published the result in the Quarterly Journal of Economics in 2025. Issues resolved per hour rose 15% on average, and 34% among the newest and least experienced. The most experienced gained almost nothing. Customers were happier, and retention went up.

The common version says support teams shrink as chatbots take the first reply. I had been saying it too, without a source. Here the people stayed and the work got easier. The deflection rates quoted for the other version circulate through vendor blogs with no study behind them.

The people helped most by a machine that answers are the people with the least to judge the answer by.

Neither paper cites the other. Putting them together is mine, and I am marking it as an argument. It is also the plainest reason I know to care what a system is allowed to say. The person who needs the guardrail is the person least likely to notice it is missing.

The proof stays behind

Avinash Bhat, Ian Arawjo, Disha Shrivastava and Jin Guo interviewed 31 experienced technical writers in 2026 about reviewing documentation with these tools. Internal systems, ticket trackers, wikis, customer data: none of it can go to an outside provider, so writers paste in fragments and the structure stays behind. A majority, 25 of the 31, said the models give feedback without enough context, including telling them a document was complete when the model had no way to know. One called them very people pleasing. The writer stayed responsible for whatever came out.

The answer carries fewer of the cues a person judges by, it reaches someone more likely to be new to the subject, and it was assembled from material that lost its structure on the way in.

That is the case for grounding a system in your own material, and the reason it is a design job. Somebody chooses which sources it may quote, puts them where it can reach them, and makes it show which one an answer came from.

Where each job sits

The question I was asked was about position, so I drew the position. This is who writes what, and where you meet it.

engineering product management conversation design ux design ui design ux research how do I get access? SOMEONE ASKS ux writing guides you through the screen you are on technical documentation explains how it works, and how to fix it when it breaks chatbot where a model speaks as your product. What it draws on is a decision service design support works from the same documents, and now sits behind a chatbot that answers first marketing and sales brings you to the product, and is measured on whether you arrive forums and other users other people answer, and nobody at the company can correct it general models everyone here works through this layer, and nobody here owns it data counts what you did at every layer, and never records why you did it

Where the answers in your product come from

the four I do, darker on the surfaces you meet everyone else in the product two layers that count nobody
click the map to open it full size
research goes out to the person and comes back as evidence.
data makes the return trip, in green, and counts you at every layer it crosses.
two routes out of that band. One goes straight to the person, crossing everything your team wrote. The other is wired into your own chatbot and answers in your name from what it read elsewhere, which is what happens when a product ships a chat with nothing of its own behind it.
support works from the same documents and now sits behind a chatbot that replies first.
the inner hatched band is service design, which arranges the people and processes behind everything else and writes nothing itself.
the outer hatched band is general models, sitting on the edge of the company: a layer everyone here works through and nobody here owns.
a short line ties a name to the slice it belongs to.
Angle carries one meaning in every ring: how much of the answering that surface does. Every ring is divided, so none of them reads as all of something, and thickness is how many people work in the layer. Conversation design sits directly outside the chatbot because that is the surface it designs. The proportions are my estimate and nothing measures them. Everything named in the key is sourced at the end of the piece.

RAG fixes the wrong half

We use retrieval, and the bot cites its sources. That is the answer I get most often.

Retrieval solves the model’s problem. It gives the system something true to say. The person is left with a different one, which is telling whether it is true for their case. A citation shows which document a sentence came from. It shows nothing about whether that document is the one for your version, your plan, your region.

Accuracy on its own buys less than it looks like it buys.

An answer needs an address

Three things close that gap, and all three are documentation work.

So grounding does two jobs. It gives the model true sentences, and it gives the person back the map.

Before the answer

Diana Deibel and Rebecca Evanhoe (2021) put conversation design inside UX design, with a tight focus on talking, and describe practitioners who start from what people need and the words they use to ask. Erika Hall (2018) goes further, and applies conversational principles to any interface, whatever its mode.

Both rest on the same practical fact. Somebody decides in advance what a system may say and where each answer comes from, and that happens before anybody asks anything.

Plenty of products ship a chat before anybody makes those calls. No documentation for it to draw on, no conversation design over it, a general model wired to the product and released. The answers come out coherent. They also come out in nobody's voice, and what they offer as fact is whatever is generally said to be true about software like yours.

That failure is the hardest one to see, because it looks like a working feature. The chatbot replies, the sentences are clean, and nobody can tell the answer came from the open internet wearing your brand.

That is the job I was asked about, and that is where it sits. Before the answer.

I hold this the way I hold my own work. The assistant in the corner of this page is mine, and I gave it a coverage metric that could not see its worst failures. That went as well as it sounds. So read the next section before you quote me.

What I cannot say

Sources

  • Write the Docs (2026). State of Docs 2026, section Docs and product. 1,131 respondents. stateofdocs.com
  • Pirolli, P., and Card, S. K. (1999). Information Foraging. Psychological Review, 106(4), 643–675. doi.org/10.1037/0033-295X.106.4.643
  • Kim, H., Kosakovsky Pond, S. L., and MacNeil, S. (2025). Conversations over Clicks: Impact of Chatbots on Information Search in Interdisciplinary Learning. IEEE Frontiers in Education Conference (FIE), Nashville. arxiv.org/abs/2507.21490
  • Brynjolfsson, E., Li, D., and Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889–942. academic.oup.com
  • Bhat, A., Arawjo, I., Shrivastava, D., and Guo, J. L. C. (2026). “A Second Set of Eyes”: The Process and Challenges of Software Documentation Review. Preprint, accepted at ACM GROUP 2027. arxiv.org/abs/2608.26232. Participants were recruited through the Write the Docs community and the LinkedIn Technical Writer Forum.
  • Deibel, D., and Evanhoe, R. (2021). Conversations with Things: UX Design for Chat and Voice. Rosenfeld Media. rosenfeldmedia.com
  • Hall, E. (2018). Conversational Design. A Book Apart. abookapart.com
  • Gibbons, S. (2017). Service Design 101. Nielsen Norman Group. nngroup.com

If a study here is yours and I summarised it badly, tell me: write@elizamarin.com. Corrections go into the piece with a note. The four disciplines in the map are written up at conversation design, technical writing, ux writing and ux research.