← All writing

An AI response is more than a typing animation

Design the waiting, streaming, stopping, and recovery states that make an AI feature feel dependable.

The first version of an AI interface often has a text box, a send button, and an animated response. It looks convincing when the network is fast and the model finishes without interruption.

The harder design work starts when a request takes too long, the reader presses Stop, or the connection breaks halfway through an answer. A useful AI feature makes those situations understandable.

Separate waiting from receiving

Treat the start of a request and the arrival of text as different states. After Send, acknowledge the action immediately. When the first content arrives, replace the waiting state with the response.

A simple lifecycle might be:

idle → waiting → receiving → complete
                 ↘ stopped
        ↘ failed   ↘ failed

This is a product model, not a transport protocol. A backend may report additional events, such as a tool call or a queued task. Translate those into language that tells the reader what is happening.

Avoid inventing progress percentages for work whose completion time you cannot estimate. “Preparing your answer” is more honest than a bar that sits at 99 percent.

Define the event contract

Streaming needs a shared contract between the client and the server. For example, your application might define events for a text fragment, a finished response, and an error:

type AnswerEvent =
  | { type: 'text'; requestId: string; text: string }
  | { type: 'complete'; requestId: string }
  | { type: 'error'; requestId: string; message: string };

This is an illustrative application contract, not an API provider’s event format. Your backend must translate the actual provider events before the client can use it.

If you use server-sent events, parse the event framing correctly. A network chunk is not necessarily a complete message. An event can span multiple reads, and one read can contain multiple events. The MDN guide to server-sent events explains the wire format.

Make Stop a real action

Stopping should immediately change the interface. Preserve the text already received, mark the response as stopped, and let the reader send another request.

In a web client, an AbortController can cancel a fetch request and consumption of its response body. See the AbortController reference. Cancellation of the browser request does not by itself prove the model provider stopped generating: the server needs to handle disconnection and propagate cancellation where supported.

Give each request an ID. If an old callback arrives after a new request begins, ignore it rather than appending its text to the new answer.

Design recovery without losing work

A partial answer can still be useful. If the connection drops, keep the text and make the failure visible. Do not present an interrupted answer as complete.

A Retry action should explain whether it starts a new answer or resumes an existing operation. Blindly repeating a request can duplicate tool actions or other side effects. Separate generating text from performing an action, and design idempotency for operations that must not run twice.

For long answers, let the reader scroll upward without the page repeatedly pulling them to the bottom. Resume automatic scrolling only when they return near the latest text.

Measure what the reader experiences

Useful measurements include time to first content, total response duration, stopped responses, and failures after partial output. Track them separately so a fast first token does not hide a slow or unreliable finish.

A polished streaming feature is one where the reader knows what is happening, can stop it, and can recover without losing their place. The typing animation is just one small part of that experience.

Something to add? I’d love to hear it.

Let’s talk about it ↗