Generative UI

Designing for latency: making streaming LLM interfaces feel tactile

LLM tokens arrive piecemeal over seconds. If your interface relies on spinners and jittery text shifts, the experience feels broken. Here is how to engineer fluid, stable generative UI.

In traditional web development, latency is a binary state: you send a request, wait behind a spinner, and receive a complete response. If the payload takes 400 milliseconds, you optimize your database query. If it takes three seconds, you show a loading skeleton.

Generative AI invalidates that entire playbook.

When a user triggers an LLM task—generating a complex summary, querying a product database, or rendering an interactive UI element—the model does not return a finished bundle. It streams tokens. Over the course of five, ten, or thirty seconds, tiny fragments of words arrive across a Server-Sent Events (SSE) or WebSocket connection.

If your frontend treats streaming like a normal HTTP request, the experience feels frantic: paragraphs bounce up and down, scroll positions jump violently, and sudden layout reflows cause cumulative layout shift (CLS).

BlockingWaits for the whole answer
How do I stop streamed text from jumping around?

Buffer tokens as they arrive and flush them once per frame. Then reveal each word with a quick blur-to-sharp fade, so the eye follows meaning instead of jitter.

  • Reserve space before blocks stream in
  • Pin scroll only while the reader is at the bottom
  • Set text-wrap: stable so lines never reflow
SourcesMDNweb.dev
First word –·Shifts 0
NaiveAppends every chunk
How do I stop streamed text from jumping around?

Buffer tokens as they arrive and flush them once per frame. Then reveal each word with a quick blur-to-sharp fade, so the eye follows meaning instead of jitter.

  • Reserve space before blocks stream in
  • Pin scroll only while the reader is at the bottom
  • Set text-wrap: stable so lines never reflow
SourcesMDNweb.dev
First word –·Shifts 0
TactilePaced, word-by-word deblur
How do I stop streamed text from jumping around?

Buffer tokens as they arrive and flush them once per frame. Then reveal each word with a quick blur-to-sharp fade, so the eye follows meaning instead of jitter.

  • Reserve space before blocks stream in
  • Pin scroll only while the reader is at the bottom
  • Set text-wrap: stable so lines never reflow
SourcesMDNweb.dev
First word –·Shifts 0

Building tactile, calm generative interfaces requires understanding the interaction physics of streaming. Here are the core engineering principles we use when building generative UI at Blinkk.


1. Frame-rate budgeting during token streaming

A fast modern LLM can emit between 40 and 100 tokens per second. If you append incoming tokens directly to your application state on every chunk event:

TSX
// The naive approach: re-rendering on every token chunk
eventSource.onmessage = (event) => {
  setText((prev) => prev + event.data);
};

You force the browser to trigger a full React component re-render, Virtual DOM diff, and layout recalculation up to 100 times every second. On lower-powered mobile devices or complex DOM trees, this completely exhausts the main thread, causing dropped frames and sluggish scrolling.

The RAF-buffered token queue

Instead of re-rendering on every network packet, buffer incoming tokens in a lightweight queue and flush them to the DOM on the browser's native animation cadence using requestAnimationFrame:

TypeScript
class TokenStreamer {
  private queue: string[] = [];
  private isRafScheduled = false;
  private onFlush: (text: string) => void;

  constructor(onFlush: (text: string) => void) {
    this.onFlush = onFlush;
  }

  push(chunk: string) {
    this.queue.push(chunk);
    if (!this.isRafScheduled) {
      this.isRafScheduled = true;
      requestAnimationFrame(() => this.flush());
    }
  }

  private flush() {
    this.isRafScheduled = false;
    if (this.queue.length === 0) return;
    const batch = this.queue.join('');
    this.queue = [];
    this.onFlush(batch);
  }
}

By throttling DOM mutations to requestAnimationFrame, you guarantee that your rendering never exceeds 60fps (or 120fps on ProMotion displays), leaving the main thread free to handle smooth user scrolling and touch gestures.


2. Preventing Cumulative Layout Shift (CLS) on dynamic blocks

One of the most jarring aspects of generative UI is watching a container expand unpredictably down the screen.

When an LLM response includes markdown tables, code blocks, or embedded preview widgets, the element's height jumps from zero pixels to several hundred pixels mid-stream. If the user is reading text below that container, the content is abruptly pushed off their screen.

UnreservedTable claims space when data lands
Generate a quarterly performance breakdown

Here is the Q3 breakdown. Revenue grew while churn and CAC both fell:

MetricQ2Q3Change
Revenue$3.6M$4.2M+18%
Churn1.6%1.2%−0.4pt
CAC$480$420−12%
SourcesQ3 report
Shifts 0
ReservedSkeleton holds the space from the tool call
Generate a quarterly performance breakdown

Here is the Q3 breakdown. Revenue grew while churn and CAC both fell:

MetricQ2Q3Change
Revenue$3.6M$4.2M+18%
Churn1.6%1.2%−0.4pt
CAC$480$420−12%
SourcesQ3 report
Shifts 0
Predictive skeleton reservation

To eliminate layout jank:

  1. Container aspect-ratio reservation: If you know an incoming generation will contain a chart or data table (via function calling or tool metadata), render a container with an explicit min-height and a subtle pulse shimmer before the first token of that table arrives.
  2. Smooth height interpolation: Use modern CSS transitions with interpolate-size: allow-keywords and @starting-style so newly inserted structural blocks expand with fluid ease rather than snapping instantly.

3. Intelligent scroll anchoring

When generating long-form answers, users naturally split into two behaviors:

  • The Spectator: Wants to watch the cursor generate and stay pinned to the very bottom of the stream.
  • The Reader: Scrolls up to re-read the first paragraph while the model continues generating subsequent paragraphs below.

If you blindly execute window.scrollTo({ top: document.body.scrollHeight }) on every chunk, you forcibly yank The Reader back down to the bottom, breaking their reading flow.

The user-intent scroll listener

Detect whether the user has intentionally scrolled up away from the bottom edge:

TypeScript
function shouldAutoScroll(container: HTMLElement, threshold = 64): boolean {
  const distanceFromBottom =
    container.scrollHeight - container.scrollTop - container.clientHeight;
  return distanceFromBottom <= threshold;
}

If distanceFromBottom exceeds the threshold, the user has taken manual control: disable auto-scroll immediately. Provide a floating, non-intrusive badge ("New content arriving below ↓") that allows them to re-anchor when they choose.


4. Graceful handling of network turbulence

Streaming connections are fragile. On mobile connections, switching between Wi-Fi and 5G frequently interrupts an active SSE connection.

In an unhardened prototype, a dropped stream leaves the user staring at half-generated gibberish with a frozen blinking cursor. In an enterprise system:

  • Token checkpoints: Maintain an immutable log of rendered tokens in local component state. If the connection drops, attempt an automatic reconnection with a Last-Event-ID header so the server can resume stream delivery from the exact character where it stopped.
  • Explicit error boundaries: If a stream fails irreparably, never erase what has already rendered. Display a graceful inline status: "Generation paused due to connection interruption" with a single-click "Resume generating" action.

Craft is the antidote to AI sameness

As generative AI becomes a standard primitive across modern web applications, the differentiator is no longer whether your site can talk to an LLM.

The differentiator is how it feels.

A generative interface that stutters, shifts, and breaks on slow networks communicates carelessness. An interface that buffers tokens smoothly, respects user scroll intent, and reserves layout geometry feels intentional, calm, and trustworthy.

In the AI era, performance and interaction craft matter more than ever.