Playbook · 6 minute read
How to Build an Elasticsearch AI Agent
An Elasticsearch AI agent combines BM25 keyword scoring with dense vector retrieval through hybrid queries, enforces document-level and field-level security by executing as the requesting user, generates query DSL constrained to known indices and fields, and for observability data reasons over aggregations rather than raw documents. Hybrid retrieval and security enforcement are where the value and the risk both sit.
Elasticsearch already provides the two retrieval mechanisms agents need, lexical scoring and vector similarity, in one engine with document-level security attached. It also holds the logs and metrics that operations agents reason over. That makes it a strong foundation and one where the failure modes are specific: hybrid retrieval that was never tuned, security that a service account bypasses, and log queries that bring the cluster to its knees. This guide covers building agents that avoid all three, drawing on FISTA Solutions' AI agents delivery on search and observability platforms. It complements how to build a hybrid search system and how to build a log analysis agent.
Why hybrid retrieval?
Because enterprise queries are mixed. A user asking about an error in a specific service needs the exact error string matched lexically and the surrounding meaning matched semantically. BM25 handles product codes, identifiers, names, and exact phrases; dense vectors handle paraphrase and intent. Either alone misses half the queries.
| Retrieval mode | Strong on | Weak on |
|---|---|---|
| BM25 | Identifiers, exact phrases, rare terms | Paraphrase, intent |
| Dense vector | Meaning, paraphrase | Exact identifiers, rare terms |
| Hybrid with rank fusion | Both | Requires tuning per corpus |
| Hybrid with reranking | Both, higher precision | Latency and cost |
Elasticsearch supports both in one query, with reciprocal rank fusion or weighted combination, and a reranking stage where precision justifies the latency. The tuning is per corpus: the right balance for support articles differs from the right balance for contracts. See what is hybrid search.
How is security enforced?
By executing as the requesting user. Document-level security restricts which documents a role can see; field-level security restricts which fields. Both attach to roles and apply when the query runs under an identity holding those roles, through per-user API keys or token-based impersonation.
An agent running as a broad service account bypasses both. It returns documents its users should not see, and because retrieval ranks across the whole index, it leaks even when results are filtered afterwards. The permission test, querying as two users with different document-level security and confirming results differ, comes before launch. See ai access control.
How should query generation be constrained?
To known indices and fields from the mappings, a safe subset of query DSL, mandatory size limits, and timeouts. Query DSL is expressive enough to include scripts, regular expressions, wildcards on leading characters, and deep aggregations, all of which can be catastrophically expensive. The agent's generated queries should exclude those constructs entirely.
Mappings are the grounding material, and they benefit from the same enrichment as any schema: field descriptions, notes on which fields are analysed versus keyword, and which are indexed for vector search. The generated query is returned with results so engineers can inspect and reuse it.
How do agents work over logs and metrics?
By aggregating first. Log indices hold millions of documents per day, and a model cannot read them. What it can read is a terms aggregation showing the top error types, a date histogram showing when a spike began, a percentile aggregation showing latency distribution, or a cardinality estimate of affected hosts.
The agent's pattern is therefore: aggregate to find the shape of the problem, narrow to the specific pattern, then fetch a small sample of raw documents for that pattern only. Fetching raw logs first is both useless to the model and dangerous to the cluster. See how to build a log analysis agent.
How is the cluster protected?
With a dedicated role and explicit limits. The agent gets its own role with document-level security where relevant and no write access to operational indices. Search timeouts bound query duration. Hard limits cap result size and aggregation bucket counts. The agent's queries are tagged so their load is visible separately in monitoring, and circuit breaker trips attributable to the agent trigger review.
On clusters shared with production search and ingestion, this matters most: an agent's exploratory aggregation over a year of logs can affect every user of the cluster.
What about ingestion pipelines?
Agents that write to Elasticsearch, such as enriching documents with classifications or embeddings, should do so through ingest pipelines or bulk operations with idempotency, under a role scoped to the target index, and never to operational indices that serve production search. Enrichment at ingest time, where a pipeline calls out for classification or embedding, keeps derived fields consistent and avoids separate reindexing.
How is it evaluated?
For retrieval, against real queries with known relevant documents, measuring recall at k and reranked precision, tuned per corpus. For security, as multiple users with different document-level rules. For log agents, against incidents with known root causes, measuring whether the agent's aggregation sequence found the pattern and how many raw documents it fetched to do so.
What does the build sequence look like?
One week on roles, per-user execution, and the permission test. One to two weeks on mapping documentation for the indices in scope. Two weeks on hybrid retrieval tuning against a reference query set. One week on constrained DSL generation. For log agents, two weeks on aggregation strategies with the operations team supplying incident cases.
What goes wrong?
Service accounts bypassing document-level security. Vector-only retrieval that misses identifiers. Unconstrained DSL with scripts and leading wildcards. Log agents fetching raw documents. No size or timeout limits. Agent load indistinguishable from user load in monitoring. And hybrid retrieval deployed with default weights that were never tuned for the corpus.
How should the reference query set be built?
From search logs, not from imagination. Elasticsearch already records what people searched for, and queries that returned nothing or that were immediately refined are the ones the agent must handle. Pair each with the documents a domain expert judges relevant, include the identifier-heavy queries that vector search struggles with and the paraphrased queries that BM25 struggles with, and the set will exercise the hybrid balance properly. Re-run it whenever the corpus, the mappings, or the embedding model changes, because retrieval tuning drifts silently.
How FISTA Solutions helps
FISTA Solutions builds Elasticsearch agents with tuned hybrid retrieval, per-user execution so document-level security holds, constrained DSL generation with inspectable queries, aggregation-first log reasoning, and cluster protection through dedicated roles and limits, through AI enablement, AI agents, and forward deployed engineers working with search and platform teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To build retrieval and log agents that respect the cluster and its permissions, message FISTA on WhatsApp, or read how to build a hybrid search system.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why use hybrid retrieval in Elasticsearch?
Because enterprise queries mix semantic intent with exact identifiers such as product codes, error strings, and names, and BM25 handles the exact matches while vector search handles the meaning. Combining them with reciprocal rank fusion or weighted scoring outperforms either alone on real corpora.
02How is security enforced?
Through document-level and field-level security attached to roles, which apply when the agent executes queries as the requesting user via API keys or token-based impersonation. An agent running as a broad service account bypasses both and returns documents its users should not see.
03How should query generation be constrained?
To a documented set of indices and fields drawn from mappings, a safe subset of query DSL that excludes scripts and expensive constructs, mandatory size limits, and timeouts, with the generated query returned alongside results so engineers can inspect it.
04How do agents work over log data?
By reasoning over aggregations: counts, terms, histograms, and percentiles across time windows, which summarise millions of documents into a shape a model can interpret, then fetching a small sample of raw documents only for the specific pattern identified.
05How is the cluster protected?
With a dedicated role for the agent, search timeouts, hard limits on result size and aggregation cardinality, circuit breaker awareness, and monitoring of the agent's query load separately, so an enthusiastic agent cannot degrade search or ingestion for everyone else.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.