Web & Mobile ┬╖ 5 minute read
WebSocket Architecture: Connections, Scale and Failure
WebSockets give bidirectional real-time delivery at the cost of stateful connections, which complicates scaling, deployment, and failure handling. Most applications that reach for them need only server-to-client push, which simpler transports handle with less operational cost.
WebSockets give bidirectional real-time delivery and introduce stateful connections, which complicate scaling, deployment, and failure handling. Most applications reaching for them need less. This guide covers the trade, drawing on FISTA Solutions' web and mobile work.
What are the options for real-time delivery?
Several, and the simplest that fits is usually right.
| Transport | Suits |
|---|---|
| Polling | Low frequency, simple, stateless |
| Long polling | Moderate frequency, wide compatibility |
| Server-sent events | Server-to-client push, simple, reconnects natively |
| WebSockets | Bidirectional, frequent client sends |
| WebTransport | Newer, unreliable-datagram cases |
| Push notifications | Delivery when the app is closed |
When do WebSockets earn their cost?
When the client sends frequently as well as receiving: collaborative editing, multiplayer interaction, live control surfaces, and chat with typing indicators.
For server-to-client updates alone тАФ notifications, live figures, progress тАФ server-sent events give you the same delivery with automatic reconnection, simpler infrastructure, and no protocol upgrade.
Teams frequently choose WebSockets by default and pay the operational cost for capability they do not use.
How does scaling work?
Connections are stateful and live on one instance, which changes several things at once.
Load balancing must distribute connections rather than requests, and a balancer that rebalances aggressively disconnects users. Broadcasting requires a broker between instances. Instance capacity is measured in concurrent connections rather than requests per second.
Plan the broker from the start. In-memory fan-out works perfectly on one instance and fails silently the moment a second exists, which is a difficult bug to find.
What should reconnection do?
Reconnect automatically with exponential backoff and jitter, recover missed state, and tell the user only when it matters.
Networks drop connections constantly, particularly on mobile. An application that reconnects transparently and catches up feels reliable; one that requires a refresh feels broken even though the failure was the network.
State recovery is the harder half. Sequence numbers with replay on reconnect, or a resync endpoint, are the usual approaches, and both need designing rather than adding after users report missing messages.
How do deployments work?
Carefully. Restarting an instance disconnects every client on it, which produces a reconnection storm.
Drain connections gracefully: stop accepting new ones, notify clients to reconnect over a spread interval, then shut down. Clients reconnecting simultaneously after a deploy can overwhelm the remaining capacity.
This is the operational cost people underestimate. Stateless services deploy without anyone noticing; connection-oriented ones need a procedure.
What about authentication and authorisation?
Authenticate at connection time and re-check authorisation for actions, not only at the handshake.
A connection established an hour ago carries whatever permissions were valid then. Long-lived connections need either periodic revalidation or per-message authorisation, depending on how quickly permissions change in your system.
Token expiry during a connection is the case teams miss. Decide whether to disconnect, refresh, or degrade.
How do you handle backpressure?
By bounding per-connection queues and deciding what to drop.
A slow client that cannot consume messages as fast as you produce them will accumulate a queue. Unbounded queues consume memory until the instance fails; bounded queues require a decision about what to discard.
For state updates, dropping intermediate values and sending the latest is usually right. For events that each matter, disconnecting the client and letting it resync is better than silently losing messages. See what is backpressure in AI pipelines.
What are the common mistakes?
Choosing WebSockets for one-directional push. In-memory fan-out. No reconnection strategy. Deploys without connection draining. Authorisation checked only at handshake. And unbounded per-connection queues.
How do you test it?
Test reconnection by dropping connections deliberately, test fan-out across at least two instances, test deployment draining, and test slow-consumer behaviour.
Mobile network conditions are worth simulating specifically, because connection drops there are frequent and the recovery path is what users experience.
What does it cost to operate?
Connection-holding capacity, a broker for fan-out, and the operational attention that stateful services require. Higher than request-response for the same traffic.
Managed real-time services shift that cost to a per-connection fee, which is frequently cheaper than operating it yourself below substantial scale.
What should you measure?
Concurrent connections per instance, reconnection rate, message delivery latency, dropped messages, and connection duration distribution.
How does this apply to AI streaming?
Streaming AI responses is server-to-client push, which means server-sent events usually fit better than WebSockets.
The exception is a conversational interface where the client sends frequently, or where the same connection carries other real-time features. Otherwise the simpler transport gives the same experience with less to operate. See streaming UI patterns for AI apps.
When is this the wrong approach?
When updates are infrequent, when the client only receives, or when the team cannot take on stateful operational complexity. Polling at a sensible interval is unfashionable and adequate for a great deal of what teams build real-time infrastructure for.
What should you do first?
Check whether your clients actually send frequently. If they only receive, server-sent events will give you the same experience with considerably less to operate.
How FISTA Solutions helps
FISTA Solutions builds and operates production systems through web and mobile, AI enablement, and staff augmentation: transport chosen for what the application actually needs rather than by default, reconnection and state recovery designed before launch, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To scope this work, message FISTA on WhatsApp, or read streaming UI patterns for AI apps.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01When do you actually need WebSockets?
When the client needs to send frequently as well as receive тАФ collaborative editing, multiplayer interaction, live control. For server-to-client updates alone, server-sent events or polling are simpler and cheaper to operate.
02Why do stateful connections complicate scaling?
Because a connection lives on one instance. Broadcasting to all clients requires a broker between instances, load balancing must account for connection distribution, and deployments disconnect everyone unless handled carefully.
03What determines perceived reliability?
Reconnection. Networks drop connections constantly on mobile, and an application that reconnects transparently with state recovery feels reliable while one that requires a refresh does not.
04How do you handle fan-out?
With a broker or pub-sub layer between instances, so a message published anywhere reaches clients connected everywhere. In-memory fan-out works on one instance and breaks silently on the second.
05What about delivery guarantees?
They are yours to provide. WebSockets give an ordered stream while connected and nothing across a disconnection, so any at-least-once requirement needs sequence numbers and replay on reconnect.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.