Web & Mobile · 5 minute read
Push Notification Architecture: Useful Without Being Annoying
Push notifications are the fastest route to an uninstall when handled badly. The architecture that works treats tokens as expiring, asks permission at a moment the value is obvious, batches aggressively, deep links precisely, and measures uninstalls alongside opens.
Push notifications are the fastest route to an uninstall when handled badly, and one of the few reliable retention mechanisms when handled well. This guide covers the architecture that makes the difference, drawing on FISTA Solutions' web and mobile work.
What does the pipeline look like?
Five stages between an event and a device.
| Stage | Responsibility |
|---|---|
| Event | Something happened worth telling someone |
| Eligibility | Preferences, quiet hours, frequency caps |
| Aggregation | Batch related events into one message |
| Composition | Copy, destination, locale |
| Delivery | Platform services, token handling, retries |
| Feedback | Opens, dismissals, invalid tokens |
Why does permission timing dominate everything?
Because on both platforms a refusal is close to permanent.
Users who deny at first launch must go into system settings to change their mind, and almost nobody does. That single prompt determines whether the channel exists for that user at all.
Ask when the benefit is concrete: after they set up a price alert, after they place an order, after they join a conversation. Explain what you will send before the system prompt appears, so the prompt is a confirmation rather than a surprise.
How should tokens be handled?
As disposable values tied to a device, not a user.
Tokens change on reinstall, on device restore, and occasionally without an obvious cause. Refresh on every launch and update the server, or you will be sending to addresses that no longer exist.
One user may have several devices, so the mapping is one-to-many. When the platform reports a token invalid, delete it immediately — continuing to send to dead tokens degrades your sending reputation and wastes quota.
Where does batching belong?
In the pipeline, so every feature inherits it.
If each feature decides its own sending, nobody owns the total. A user receiving three notifications from three features in five minutes experiences one annoying app, not three reasonable features.
Apply a frequency cap per user across all senders, aggregate related events into one message, and hold non-urgent notifications for a delivery window. The aggregate is what the user judges.
How should deep linking work?
Precisely, with the navigation stack constructed so onward movement makes sense.
A notification about a specific message should open that message. A user who lands on a home screen and has to search wasted their tap, and learns to ignore the next notification.
Handle the logged-out and cold-start cases: the destination must survive authentication and app launch. See mobile app architecture guide.
What preferences should users have?
Categories, not a single switch.
A user who wants order updates but not marketing should be able to say so. An all-or-nothing control means the marketing message costs you the transactional channel too.
Quiet hours in the user's own timezone are the other essential control. A notification at three in the morning is remembered, and not favourably.
How do you make delivery reliable?
By treating the platform services as best-effort and designing around it.
Notifications can be delayed, throttled, or dropped, particularly on low-priority channels or when a device is in a power-saving state. Anything important must also be visible in the app.
Use high-priority delivery only for genuinely time-sensitive messages. Abusing it gets an application throttled by the platform, which degrades everything you send.
What are the common mistakes?
Asking for permission at first launch. Per-feature sending with no global cap. Notifications that open the home screen. A single on-off preference. Stale tokens. And measuring opens without watching uninstalls.
How do you test it?
Test cold start from a notification, logged-out deep links, permission denial paths, and behaviour when the app is in the foreground.
Test with the device in a power-saving state, because that is where delivery assumptions break.
What does it cost to operate?
Platform delivery is free; the cost is the pipeline, the preference system, and the engineering to keep token handling correct.
A managed messaging service shortens this considerably and is usually the right call unless volume or requirements are unusual.
What should you measure?
Permission grant rate by prompt context, delivery rate, open rate by category, notification disable rate, and uninstall rate in the days following a campaign.
Can AI improve notification quality?
For timing and relevance, yes — a model can predict when a given user is likely to engage, and suppress messages likely to be ignored.
Keep generated copy under review before it sends. A notification is a direct, unfiltered message to a user, and an unreviewed generated one is a brand risk with no recall mechanism. See human in the loop AI explained.
When is this the wrong approach?
An application people open deliberately several times a day does not need push notifications to drive return visits, and adding them is more likely to cost engagement than gain it.
What should you do first?
Look at your permission prompt. If it appears at first launch, moving it to a moment of obvious value is the single highest-return change available.
How FISTA Solutions helps
FISTA Solutions builds and operates production systems through web and mobile, AI enablement, and staff augmentation: a global frequency cap in the pipeline rather than per-feature sending, and permission prompts placed at a moment the value is already visible, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To scope this work, message FISTA on WhatsApp, or read mobile app architecture guide.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01When should you ask for notification permission?
At the moment the user would obviously benefit — after setting up an alert, placing an order, or joining a conversation. Asking at first launch, before any value is visible, produces a refusal you rarely get to reverse.
02How should device tokens be managed?
As expiring values that change on reinstall, restore, and sometimes at random. Refresh them on every launch, store them per device rather than per user, and delete them when the platform reports them invalid.
03Why does batching matter?
Because ten separate notifications for related events is what makes people disable them. Aggregation belongs in the pipeline so every feature benefits, not in each feature's own logic.
04What should a notification link to?
The exact item it concerns, with the back stack constructed so the user can navigate onward. Dropping someone on a home screen wastes the tap and trains them to ignore the next one.
05What should you measure?
Open rate matters less than the disable rate and the uninstall rate in the days after a campaign. A notification with strong opens that raises uninstalls is a net loss you will not see otherwise.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.