FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook · 5 minute read

How to Build a Neo4j AI Agent

A Neo4j AI agent grounds in the graph schema of labels, relationship types, and properties, generates Cypher constrained to read-only patterns with depth and result limits, combines vector indexes for semantic entry points with traversal for multi-hop questions, and runs under roles that scope which labels and relationships it may traverse. Multi-hop reasoning is where graphs beat other stores.

By FISTA Solutions· AI-Native Engineering Team·
How to Build a Neo4j AI Agent article cover

Some questions are about connection rather than aggregation: which suppliers depend on the same sub-tier factory, how a customer entity links to a sanctioned party through intermediaries, which systems a change to one component affects, who worked with whom through which projects. Relational stores answer those with joins that multiply, and vector stores do not answer them at all. Graph databases answer them natively, and an agent on Neo4j can exploit that if it generates Cypher that is grounded, bounded, and permissioned. This guide covers building one, drawing on FISTA Solutions' AI agents delivery on graph and knowledge systems. It complements how to build a graph rag system and what is a knowledge graph.

What grounds the agent?

The schema, introspected and described. Neo4j exposes labels, relationship types, their directions, and property keys, and the agent needs all of them plus descriptions of what each means in the domain: that a SUPPLIES relationship runs from supplier to customer, that an ENTITY node may represent a person or an organisation distinguished by a property, that a REPORTS_TO chain represents management hierarchy.

Schema elementIntrospection providesDescription adds
Node labelsNames and countsDomain meaning, distinguishing properties
Relationship typesNames, directions, endpointsSemantics, cardinality expectations
PropertiesKeys and typesMeaning, units, freshness
IndexesWhich properties are indexedWhich lookups are efficient
ConstraintsUniqueness and existenceWhich properties identify a node

Without descriptions the agent generates Cypher using relationships that sound right and do not exist, or traverses in the wrong direction, both of which return empty or wrong results without an error.

How should Cypher generation be constrained?

Strictly, because Cypher can express queries that never finish. Constraints that hold: read-only clauses only, with write clauses rejected before execution; a maximum depth on variable-length paths, since an unbounded traversal on a connected graph is a full scan; mandatory LIMIT clauses; query timeouts; and labels and relationship types validated against the schema before the query runs.

The generated Cypher is returned with every result, along with the paths found, so an engineer can inspect, re-run, and verify. Clarification is asked when a question could map to several relationship types, which in a rich graph is common.

How does graph retrieval combine with vectors?

As entry point plus traversal. Vector indexes on node text properties find the nodes semantically relevant to a question; traversal from those nodes follows typed relationships to gather the connected context: the supplier's sub-tier suppliers, the entity's linked parties, the component's downstream dependents.

The combination answers multi-hop questions that pure vector retrieval cannot, because the relevant context is connected to rather than similar to the query, and that pure traversal cannot, because it needs somewhere to start. Results carry the path, which is the explanation. See how to build a graph rag system and what is graph rag.

How is access controlled?

Through Neo4j's fine-grained privileges, which distinguish between reading a node's properties and traversing through it. A role can permit traversal of a label without reading its properties, or reading some properties but not others, or following certain relationship types and not others.

That granularity matters for agents. An agent answering supply chain questions might traverse company nodes and SUPPLIES relationships without reading financial properties or following OWNS relationships into ownership structures it has no business seeing. Enforcing that in the database rather than in the prompt is what makes it reliable. See ai access control.

What questions suit a graph agent?

Questions about connection, path, and structure. Dependency and impact analysis. Relationship discovery between entities. Path finding with constraints. Community and cluster identification. Provenance and lineage chains.

Questions about aggregates over large populations usually suit a warehouse better, and an agent that routes those to the graph produces slow, awkward Cypher. Where an organisation runs both, a routing layer that sends connection questions to the graph and aggregate questions to the warehouse serves users better than forcing either store to do the other's work.

How is performance protected?

By bounding traversal and using indexes. Depth limits and result limits are the primary controls. Starting traversals from indexed properties rather than label scans is the second. Timeouts are the backstop. Monitoring the agent's query patterns against the query log identifies patterns that need a new index or a tighter constraint.

Unbounded traversals are the specific risk: a shortest-path query between two nodes with no depth limit on a dense graph can run indefinitely, and the agent framework must reject that shape rather than trusting the model to avoid it.

How is it evaluated?

Against real questions with the correct answer and the correct path. Measure whether the generated Cypher used the right labels and relationships, whether the path returned is the right explanation, and execution time relative to a developer-written query. Include questions that require multi-hop traversal, because they are the reason the graph exists, and questions that should be routed elsewhere, to test the routing.

What does the build sequence look like?

One to two weeks on schema introspection and description enrichment, which is the step that determines whether generated Cypher is real. One week on roles and privileges. Two weeks on constrained Cypher generation with graph engineers testing. One week on vector index entry points if semantic questions are in scope. Then routing between graph and other stores where both exist.

What goes wrong?

Cypher invented from label names without schema grounding. Unbounded variable-length paths. Write clauses reaching the database. Roles that permit reading everything reachable. Aggregate questions forced through the graph. And answers without paths, which discard the graph's main advantage: that the answer explains itself.

How FISTA Solutions helps

FISTA Solutions builds Neo4j agents grounded in described schemas, generating read-only bounded Cypher with paths returned for verification, combining vector entry points with typed traversal for multi-hop questions, and running under fine-grained privileges that scope what may be traversed, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To answer the connection questions other stores cannot, message FISTA on WhatsApp, or read how to build a graph rag system.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What does an agent need to understand a Neo4j graph?

The schema: node labels, relationship types, their directions, and the properties on each, obtained through schema introspection and enriched with descriptions of what each label and relationship means in the domain. Cypher generated without that grounding invents relationships that do not exist.

02How should Cypher generation be constrained?

To read-only clauses, a maximum traversal depth, mandatory result limits, query timeouts, and the schema's known labels and relationship types. Unbounded variable-length paths and write clauses should be rejected before execution, and the generated Cypher returned with results.

03How does graph retrieval combine with vector search?

Vector indexes on node properties find semantically relevant entry nodes for a question; traversal from those nodes follows relationships to gather connected context. The combination answers multi-hop questions that neither pure vector retrieval nor pure traversal handles well.

04How is access controlled?

Through Neo4j roles with fine-grained privileges that grant traverse and read rights on specific labels, relationship types, and properties, so an agent's role determines not just which nodes it may read but which relationships it may follow, enforced by the database rather than the application.

05What questions suit a graph agent?

Ones about connection and path: which suppliers share a sub-tier dependency, how an entity links to a sanctioned party, which components a change affects downstream, or who collaborated with whom through what. Questions about aggregate counts usually suit a warehouse better.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project