FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Comparison ¡ 4 minute read

Monolithic vs Modular AI Architecture: Where to Draw Lines

Modularity is not free, and AI systems have a few boundaries worth drawing regardless: the model access layer, the retrieval pipeline, and evaluation. Beyond those, separate only where teams, scaling, or deployment cadence genuinely differ, because coordination costs are paid continuously.

By FISTA Solutions¡ AI-Native Engineering Team¡
Monolithic vs Modular AI Architecture: Where to Draw Lines article cover

Modularity is not free, and AI systems have a few boundaries worth drawing. This guide covers which ones pay, drawing on FISTA Solutions' AI agents production work.

Which boundaries pay?

Three that almost always do, and the rest that depend.

BoundaryWorth drawing whenCost if drawn early
Model access layerAlwaysMinimal
Retrieval pipelineNon-trivial ingestionLow
EvaluationAlwaysMinimal
Agent and toolsSeparate teams or cadenceCoordination
Per-feature servicesGenuine scaling differenceCoordination
Per-model servicesRarelyCoordination for nothing

Why is the model access layer always a boundary?

Because it makes the provider replaceable at almost no cost.

A thin layer wrapping provider calls, handling retries, logging versions, and attributing cost gives you routing, failover, and migration ability later. It is a few hundred lines and it pays for itself the first time a provider changes.

It does not need to be a separate service — a module is sufficient — but the boundary should exist. See the shift from model choice to system design.

When should retrieval be separate?

Once it has its own lifecycle.

Ingestion, chunking, embedding, indexing, and refresh form a pipeline with different scaling and different failure modes from request handling. Separating them lets each be operated appropriately.

It also lets several applications share one corpus, which avoids duplicating ingestion. Below that — a small static corpus loaded at startup — separation adds nothing. See RAG quality checklist.

Why must evaluation be independent?

Because attribution requires testing components separately.

An evaluation harness that can only exercise the whole system tells you the answer was wrong. One that can test retrieval alone tells you whether the right passage was found, which is a different fix.

That requires components to be callable independently, which is an architectural constraint worth accepting early. See how to build an agent evaluation harness.

What are legitimate reasons to separate further?

Teams, scaling, and deployment cadence.

If two teams own two capabilities, a boundary matches the organisation and prevents shared ownership of everything. If one component needs to scale independently, separation lets it. If one deploys daily and another monthly, coupling them slows the fast one.

Those are concrete and observable. Anticipated future need is not.

What does coordination actually cost?

Interface maintenance, version compatibility, distributed debugging, and deployment ordering.

Every boundary is an interface that must stay compatible as both sides change. Debugging spans services. Deployments must consider ordering. Each is small; together they are the reason premature modularity slows teams.

That cost is paid every week, whereas the benefit of a speculative boundary may never arrive.

How do you introduce boundaries later?

By keeping modules clean inside the monolith first.

A well-structured single service with clear internal modules can be split when a reason appears. A tangled one cannot, which is why internal structure matters even without service boundaries.

That is the practical sequence: clean modules now, services when a concrete reason arrives. See microservices vs monolith.

How do you run your own comparison?

Ask, for each proposed boundary, what concrete problem it solves today. If the answer is about future scale or tidiness, do not draw it.

Then check whether your evaluation can test retrieval independently of generation. If it cannot, that boundary is missing and it is the one that matters most.

What does switching cost later?

Splitting a clean monolith is straightforward; splitting a tangled one is a rewrite. Merging services is usually easy and rarely done, because the boundaries have become organisational.

Internal module discipline is what keeps the option open in both directions.

What do people get wrong here?

Separating for tidiness. Anticipating scale that never arrives. Evaluation that can only test end to end. Provider calls scattered through the codebase. And tangled internal structure that makes later splitting impossible.

Does the agent change this?

It adds one consideration: tools should be independently callable and testable, which is a boundary regardless of deployment shape.

A tool implemented as a function the agent calls is easier to test and to reuse than one embedded in the agent's flow. Whether it becomes a separate service depends on the usual reasons. See fine-grained vs coarse tool design.

Which should you choose?

Draw three boundaries early: model access, retrieval once it has a lifecycle, and evaluation independence. Start monolithic otherwise, keep internal modules clean, and separate further only when teams, scaling, or cadence give you a concrete reason.

What should you do first?

Check whether your evaluation can test retrieval without running generation. If not, that is the boundary to introduce first.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: model access, retrieval, and evaluation separated early with everything else kept in one clean service until a concrete reason to split appears, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To run this comparison against your own workload, message FISTA on WhatsApp, or read the shift from model choice to system design.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Which boundaries always pay?

The model access layer, so providers are replaceable; the retrieval pipeline, once it involves real ingestion and indexing; and evaluation, which must be able to test components independently.

02Why start monolithic?

Because a single service is simpler to build, deploy, debug, and evaluate. Boundaries should be introduced when a specific problem justifies them, not in anticipation.

03When should retrieval be separate?

When it has its own ingestion pipeline, index lifecycle, and scaling characteristics — which is most production systems. It also lets several applications share one corpus.

04Why must evaluation be independent?

Because you need to test retrieval separately from generation to attribute failures. An evaluation that can only test the whole system end to end cannot tell you which half is wrong.

05What are bad reasons to separate?

Conceptual tidiness, anticipated future scale, and following a reference architecture. Each adds coordination cost immediately for a benefit that may never arrive.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project