Playbook ¡ 6 minute read
How to Build a Databricks AI Agent
A Databricks AI agent grounds in Unity Catalog metadata for tables, columns, and lineage, executes queries through SQL warehouses with the requesting user's identity so catalog permissions apply, uses vector search and model serving for retrieval and inference inside the governance boundary, and is evaluated with the platform's tooling. Unity Catalog is what makes it defensible.
Databricks combines the data, the governance, the retrieval infrastructure, and the model serving in one platform, which makes it an unusually complete place to build an agent. The centre of that is Unity Catalog: an agent that accesses data through it inherits permissions, row and column security, lineage, and audit logging, and an agent that bypasses it forfeits all four. This guide covers building agents that stay inside the boundary, drawing on FISTA Solutions' AI agents delivery on lakehouse platforms. It complements how to build a snowflake ai agent and how to build a text to sql agent.
What does Unity Catalog provide?
| Capability | What it gives the agent |
|---|---|
| Permissions on catalogs, schemas, tables | Access control without reimplementation |
| Row filters and column masks | Row-level and field-level security applied at query |
| Metadata and comments | Grounding for table and column selection |
| Lineage | Provenance from any result back to source |
| Audit logs | Every access recorded where security already looks |
| Vector search indexes | Retrieval under the same governance |
| Model registry | Governed model versions with lineage |
Everything the agent does should go through the catalog. Direct access to storage, bypassing the catalog, loses every one of these and is the failure to design against.
How does query execution work?
Through SQL warehouses, using the Statement Execution API or a SQL connector. The agent constructs SQL against catalog tables, executes with the requesting user's identity, and receives results with row filters and column masks applied by the catalog.
Identity is the critical detail. A service principal with broad grants executing on everyone's behalf sees everything; an agent executing as the requesting user, through on-behalf-of tokens or user-scoped credentials, sees what that user may see. The permission test, running as users with different grants and confirming results differ, comes before anything goes live.
What grounds the agent's SQL?
Catalog metadata: table and column names, comments, data types, primary and foreign key declarations where present, and lineage showing where data came from. Comments are the decisive input, and most lakehouses have few of them.
Enriching comments for the tables in scope, stating what each column means, its grain, and its freshness, is the preparation that determines whether the agent chooses the right table. Lineage helps too: an agent that can see a table is derived from a certified source can prefer it over an undocumented copy. See how to build a text to sql agent for the governed-query pattern.
How is retrieval handled?
Through vector search indexes built over catalog tables, which inherit the table's permissions. That means retrieval results respect the same access controls as direct queries, the embedded content stays inside the governance boundary, and lineage extends from the retrieved passage back to its source row.
Indexes should be built over curated content with ownership and freshness, following the same discipline as any retrieval system. Indexing every table produces retrieval that surfaces stale and contradictory material. See the enterprise RAG reference architecture whitepaper.
Where does inference run?
Through model serving, either for models the organisation hosts or through governed external model endpoints that route to a provider under the platform's controls. That keeps the model call, the data it saw, and the logs in one boundary, with rate limits and cost tracking applied centrally.
It also makes the model a configuration choice. Swapping models becomes an endpoint change with re-evaluation rather than an application change, which matters over the life of an agent. See the model routing and cost control whitepaper.
How is the agent evaluated and monitored?
With the platform's own tooling. Traces are logged to lakehouse tables where they can be queried like any data; reference sets live in tables; scoring runs as jobs; and dashboards over the evaluation results use the same tools the data team already uses. That is a genuine advantage over bolting evaluation infrastructure onto a platform that has none.
The reference set discipline is unchanged: real questions, verified answers, separate retrieval and generation scoring, and continuous runs on change. See the RAG evaluation methodology whitepaper.
How is cost controlled?
Through warehouse sizing and auto-stop, query timeouts and row limits, model serving rate limits, and cost attribution by user and use case using the platform's usage tables. An agent is a workload and should have a budget and a monitored cost per question from the start.
What should the agent do first?
Governed question answering over well-documented tables for a single domain, with the permission test passed and provenance returned with every answer. Then retrieval over curated content. Then multi-step analysis. Write operations, such as an agent that creates tables or triggers jobs, only with explicit scoping and approval, since the blast radius of a write in a lakehouse is large.
What does the build sequence look like?
Two weeks on catalog comment enrichment for the domain in scope. One week on warehouse access with user identity and the permission test. Two weeks on question translation with the data team testing. One week on vector search over curated content if retrieval is in scope. Then evaluation jobs and monitoring dashboards, which are quick to build on the platform.
What goes wrong?
Storage access that bypasses the catalog. Service principals with broad grants. Sparse comments. Indexes over everything. Model calls outside the serving boundary, which lose logging and cost control. Write access granted early. And evaluation left to spreadsheets when the platform provides tables and jobs.
How does this differ from building outside the platform?
Mainly in what you do not have to build. Outside the platform, an agent needs its own permission layer mirroring the warehouse, its own retrieval infrastructure with permission synchronisation, its own model gateway with logging, and its own evaluation storage. On Databricks each of those is a platform capability the agent consumes, which shortens the build and, more importantly, keeps governance in one place security already monitors.
The trade-off is platform dependence, which is worth naming honestly. Prompts, reference sets, and evaluation logic should still be held in the organisation's own repositories so they travel if the platform decision changes; what stays platform-specific is the execution path, which is the part that should be.
How FISTA Solutions helps
FISTA Solutions builds Databricks agents that stay inside Unity Catalog, executing as the requesting user, grounded in enriched catalog metadata, retrieving through governed vector search, serving models within the boundary, and evaluated with platform tooling, through AI enablement, AI agents, and forward deployed engineers working with data platform teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To build agents on the lakehouse without leaving its governance, message FISTA on WhatsApp, or read how to build a text to sql agent.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How does an agent access Databricks data?
Through SQL warehouses via the Statement Execution API or SQL connectors, with Unity Catalog governing which catalogs, schemas, tables, and rows the identity may read. The agent grounds in catalog metadata to choose tables and columns and executes with the requesting user's identity.
02Why does Unity Catalog matter for agents?
Because it centralises permissions, row and column security, lineage, and metadata across the lakehouse, so an agent that goes through it inherits governance rather than reimplementing it, and every access is audited in one place that security teams already monitor.
03How is retrieval handled?
Through Databricks vector search over indexed tables, which sit inside Unity Catalog and inherit its permissions, so retrieval results respect the same access controls as direct queries and the embedded content never leaves the governance boundary.
04Where should the model run?
Through model serving for models the organisation hosts or through governed external model endpoints, so inference, data, and logging stay within one security boundary and the choice of model becomes a configuration decision rather than an architectural one.
05How is the agent evaluated?
With the platform's agent evaluation and tracing capabilities, logging every interaction to lakehouse tables where it can be queried, scored against reference sets, and monitored over time using the same tools the data team already uses.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.