Strategy
Spec-Driven Development: Directing AI Reliably
AI is only as reliable as the specification directing it. How spec-driven development turns probabilistic models into predictable, verifiable production systems.
FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.
FISTA field notes
Page 85 of 103.
Archive
2468 field notes · page 85 of 103
Strategy
AI is only as reliable as the specification directing it. How spec-driven development turns probabilistic models into predictable, verifiable production systems.
Playbook
Contentful separates content from presentation through a content model, which gives an agent structured fields to work with and a publishing workflow it must respect. This guide covers the Management and Delivery APIs, grounding in the content model, drafting entries and localisations into the editorial workflow, quality and SEO checks, and why the agent creates drafts rather than publishing.
Playbook
Figma holds the design system, the product screens, and the handoff artifacts engineers build from, and an agent there can catch inconsistencies no reviewer has time to find. This guide covers REST API and plugin access, design system compliance auditing against components and tokens, accessibility checks, handoff documentation generation, and why the agent proposes rather than edits.
Playbook
GitLab holds code, pipelines, issues, and deployment in one place, which lets an agent see a change from issue to production. This guide covers merge request review that reads the diff and the pipeline together, pipeline failure diagnosis from job logs, issue triage and linking, project and group access token scoping, and the protected branch and approval rules that keep the agent's contributions reviewable.
Playbook
Linear is opinionated about how software teams should work, and an agent that fights those opinions gets removed. This guide covers the GraphQL API and webhooks, triaging and enriching inbound issues, linking issues to code and customer context, cycle and project reporting, and the boundary around priority, assignment, and estimates that belongs to the team.
Playbook
Sentry captures every error but not whether anyone is looking at it, and most engineering teams have a backlog of issues nobody triaged. This guide covers building an agent that reads new issues with their stack traces, breadcrumbs, and release context, proposes root cause hypotheses, routes to the right owner, links related issues, and where appropriate drafts a fix as a pull request for review.
Playbook
WordPress runs a large share of the web, from small business sites to enterprise publishing, with a REST API and a plugin ecosystem that make agents easy to connect and easy to connect badly. This guide covers application passwords and role scoping, drafting posts and pages into the editorial workflow, content maintenance across large sites, SEO checks, and the security posture a WordPress agent requires.
Strategy
A Digital FTE is an AI system scoped to own a bounded role's work under human oversight—not a chatbot, and not a replaced person. Where it creates real leverage.
Playbook
Datadog holds the telemetry an on-call engineer needs and far more of it than anyone can read during an incident. This guide covers querying metrics, logs, and traces through the API with scoped keys, triaging alerts with correlated context, assembling incident timelines, and keeping the agent's own query load from becoming a cost line or a performance problem.
Playbook
dbt projects hold the transformation logic, lineage, tests, and documentation that an AI assistant needs to be useful to analytics engineers, in a structured form the assistant can read. This guide covers grounding in the manifest and catalogue, generating models and tests that follow project conventions, documentation generation, impact analysis from lineage, and keeping every generated change reviewable through the normal pull request flow.
Playbook
PagerDuty is where incidents become people's problems, at three in the morning, with a page and no context. An agent that briefs the responder before they open a laptop is worth more than most dashboards. This guide covers incident enrichment through the API and events, responder briefing, stakeholder update drafting, postmortem timeline assembly, and the boundary around paging, escalation, and resolution that stays with people.
Playbook
Splunk holds the machine data security and operations teams live in, queried through a language most of their colleagues never learn. An agent that turns questions into good SPL widens access and, done carelessly, exhausts search capacity. This guide covers grounded SPL generation, search head protection, role-based access enforcement, alert triage for security operations, and evaluating whether the generated searches are actually right.
Playbook
AI agents fail in the middle, wait for humans, and run for hours, which is exactly what Temporal's durable execution was built for. This guide covers structuring an agent as a Temporal workflow with model calls as activities, handling human approval through signals, retry and timeout policy for probabilistic steps, versioning prompts and logic safely, and observing what a long-running agent did.
Playbook
Infrastructure as code is the highest-blast-radius code in the organisation, and an agent that can run apply is a credential that can delete production. This guide covers what a Terraform agent should do, including plan interpretation, drift explanation, policy-checked module generation, and cost estimation, how to keep it reviewable through pull requests, and why apply authority stays with people and pipelines.
Playbook
Most enterprise AI work is batch: nightly classification, weekly embedding refresh, scheduled evaluation runs, and document backfills. Airflow already runs the data platform's batches, and it can run these too if the model calls are placed correctly. This guide covers structuring AI DAGs, pushing inference to executors rather than the scheduler, idempotent tasks, evaluation as a pipeline, and cost-aware scheduling.
Strategy
Most AI pilots impress in a demo and die before production. The real reasons—data, evaluation, governance, and adoption—and how to run one that ships.
Playbook
BigQuery is fast, serverless, and billed by bytes scanned, which means an agent that writes careless SQL is both a data exposure risk and a cost risk. This guide covers IAM and row-level security for agent identity, grounding SQL generation in dataset metadata, controlling cost through dry runs and byte limits, and returning results users can verify.
Playbook
Putting a model inside a Kafka pipeline turns every event into an inference call, which is powerful at scale and ruinous when designed carelessly. This guide covers consumer group design for inference workloads, backpressure when the model is slower than the stream, ordering and exactly-once semantics, schema governance for AI-enriched events, and keeping inference cost proportional to value.
Playbook
MongoDB's flexible schema is the reason applications choose it and the reason an agent querying it needs care: documents in one collection may not share a shape. This guide covers schema discovery, generating aggregation pipelines safely, using Atlas Vector Search for retrieval alongside operational data, scoping database roles, and keeping writes bounded.
Playbook
Graph databases answer the questions relational and vector stores struggle with: who is connected to whom, through what, and how far. An agent on Neo4j can exploit that if it generates Cypher safely and understands the graph's schema. This guide covers schema grounding, constrained Cypher generation, combining graph traversal with vector retrieval, and scoping what the agent may traverse.
Playbook
Redshift sits at the centre of many AWS data estates, with workload management queues and a cluster or serverless cost model that an agent must respect. This guide covers Data API access with IAM identity, row-level and column-level security, grounding SQL in system catalogue metadata, protecting queues from agent load, and returning results users can verify.
Playbook
Elasticsearch already does the two things retrieval agents need, keyword and vector search, and holds the logs that operations agents reason over. This guide covers hybrid retrieval combining BM25 and dense vectors, document-level security so results respect permissions, safe query DSL generation, and building agents over log and metrics data without overwhelming the cluster.
Strategy
Human-in-the-loop is not a lack of confidence in AI—it is how reliable AI systems are built. Where to place humans, and how to keep oversight without killing speed.
Strategy
Where is your organization on the AI maturity curve? The five levels—from ad-hoc tools to AI-native operations—and the one move that gets you to the next.
Start with the hard problem
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.