Research and product perspective on relational agent artifacts stored in git, state of the art (2026), and Flatbread's wedge as it pivots beyond GraphQL-only query surfaces.
The agent artifact layer in 2026 is dense with conventions and files (AGENTS.md, SKILL.md, .handoff/*.md, .GCC/branches/*, .cs/discoveries.md, vault MCPs) but thin on typed relational schemas over those artifacts. Teams get search, backlinks, versioned memory trees, and handoff packets; they rarely get reference integrity, stable cross-tool IDs, or filterable graph queries like "all blocking decisions for this effort with owning plan and producing sessions." Flatbread already models collections, refs, and rich filters over markdown/YAML in git. Positioning it as the relational layer for agent efforts in git—with MCP and generated TypeScript alongside GraphQL—is a credible category move. The recommended posture is an Proof: a preset schema (Effort → Plan, Decision, Session, Artifact, Run) plus validation, append-oriented writes, and agent-facing query APIs, without building a CMS or competing on generic databases.
Agent harnesses produce long outputs (plans, research dumps, code reviews, traces) but rarely persist them as cohesive elements bound to a named effort—feature, spike, migration, or research thread. The next phase of a pipeline often starts cold unless a human or script deliberately passes a markdown file into context.
Consequences:
- Decision drift — earlier conclusions are lost or contradicted across sessions.
- Repeated discovery — the same codebase areas get re-explained.
- Weak accountability — hard to answer "what did we decide about X for effort Y across all runs?"
- Manifest scaling — single-file hot memory (
AGENTS.md, constitution files) works until it does not; real projects scale to tiered cold stores and specialist agents (Codified Context reports on the order of tens of thousands of lines of machine-oriented specs for a ~108k LOC system).
The missing abstraction is not "more markdown" but a stable, queryable object graph whose instances live in the repo and survive tool switches.
Patterns cluster into five layers. None alone delivers typed relations + integrity over arbitrary harness layouts.
Single-file or small-set instructions loaded every session: AGENTS.md, CLAUDE.md, .cursorrules, Cursor rules with YAML frontmatter (globs, alwaysApply, description, path constraints). Empirical studies report adoption and content patterns across thousands of repos; one line of evidence associates AGENTS.md with roughly 29% lower median runtime and 17% lower output token consumption (as cited in Codified Context related work). This layer optimizes always-on priming, not structured artifact graphs.
Reusable procedures shipped as files (e.g. Claude Code Skills), Cursor Agent Skills, packaged "agentic artifacts." Agentic Beacon frames a package-manager metaphor for contexts, knowledge, and skills across teams—addressing context drift analogously to code reuse. Strength: distribution and versioning of playbooks. Weakness: still not a relational model over instances (efforts, decisions, runs).
Deliberate disk layouts for continuity: handoff-style .handoff/ trees (FEATURE.md, SPEC.md, DESIGN.md, STATE.md, SESSION.md); AgentHandoff for switching between Claude Code, Codex, Cursor with structured packets and reported 60–85% token reduction vs cold rediscovery; long-running-harness patterns (feature_list.json, progress.txt); ecosystem demand for named persistent plans (e.g. Claude Code Plan Manager discussions). This layer fixes handoff between tools or phases; it does not standardize cross-effort querying or ref validation.
Treat agent memory like version control. Git Context Controller (GCC): .GCC/main.md, per-branch commit.md, log.md, metadata.yaml; commands COMMIT, BRANCH, MERGE, CONTEXT for layered retrieval; strong benchmark results (e.g. 80.2% on SWE-Bench Verified with Claude 4 Sonnet in reported runs). Lore (arXiv) repurposes git commit trailers for structured decision shadows. agmem (GitHub): git-like, content-addressable agent memories. claude-sessions (GitHub): workspaces with discoveries, artifacts folders, checkpoints. These systems excel at timelines, branches, and checkpoints; they are weak on arbitrary typed edges (e.g. Decision → Plan → Effort) validated at index time.
Markdown vaults + retrieval: Codified Context (paper)—constitution (hot), specialist agents, cold markdown specs + MCP keyword retrieval (find_relevant_context, suggest_agent), scaled to hundreds of sessions. engraph, memory-graph, Obsidian-oriented stacks combine wiki-links, embeddings, BM25/FTS. MCP vault servers (e.g. markdown-vault-mcp, vault-mcp, vault-semantic-mcp, knowledge-mcp) expose search, backlinks, sometimes hybrid semantic + lexical retrieval. Strength: find related notes. Weakness: links are typically untyped strings; "Tier 3 document lists" are not foreign keys.
Velite, Keystatic, next-mdx-remote, and the legacy Contentlayer space optimize sites and apps from markdown collections. They overlap Flatbread on "typed-ish content in repos" but do not target multi-session agent efforts or harness-native conventions.
Across manifest, skill, workspace, memory-as-VCS, and vault layers, convergent capabilities include:
- File and folder conventions.
- Full-text and semantic search.
- Versioned or branching narrative memory.
- Wiki-links and backlinks.
Conspicuously absent as a first-class product:
| Missing capability | Why it matters |
|---|---|
| Typed relational schemas over artifacts | Filters like status, blocking, effort_id need columns, not only embeddings. |
| Reference integrity | Broken Plan → Decision links should fail at load/validate time, not at retrieval luck. |
| Stable IDs | Renames and multi-branch workflows should not orphan edges. |
| Canonical Effort (or equivalent) aggregate | Same role PR plays for commits: one object that owns the thread. |
| Non-text queries | Traverse + filter + sort without grepping or re-ranking chunks. |
Example that is painful everywhere above but natural in a relational content layer: "Decisions for Effort X where blocking: true, ordered by decided_at, with plan.title and session.tool."
Flatbread's existing architecture maps cleanly onto the gap:
- Collections and refs — GraphQL schema generation from content collections and cross-collection references (
packages/core/src/generators/schema.ts) matches Effort-linked entity graphs. - Structured filtering — Mongo-style filter operators on collection fields (root README —
eq,in,exists,regex,wildcard, etc.) exceed typical vault MCP keyword APIs for predicate-rich agent queries. - Markdown/YAML as rows —
packages/transformer-markdown,packages/transformer-yamlalign with how harnesses already emit artifacts. - Pluggable sources —
packages/source-filesystemcan ingest.agents/,.cursor/,.handoff/,.GCC/,.cs/trees as additional content paths without a new storage paradigm. - Codegen path —
packages/codegen/src/generator.tsis the right place to grow generated TypeScript accessors as the agent-ergonomic surface GraphQL is not.
The PMF audit near-term list (config typing, ID normalization, relation validation, watch mode) is the same prerequisite work an agent-artifact product needs—not a competing roadmap.
Flatbread becomes a full memory product: authoring UI, lifecycle, branching UX, primary store for traces. Target: teams wanting one vendor for agent memory. Surfaces: dashboard, editors, sync. Strength: largest narrative TAM if execution wins. Risk: collides with IDE vendors and GCC-like research stacks; violates current PMF guidance to avoid hosted CMS/dashboard; requires broad writes and permissions Flatbread does not have today.
Read-mostly preset: index existing harness folders, validate optional schemas, expose MCP + TS helpers. Target: harness engineers wiring Cursor/Claude/GCC without migrating storage. Strength: low displacement, fits today's read skew. Risk: commodity "nice indexer" unless paired with a sharp noun and integrity story.
Own the Effort aggregate explicitly: typed collections Effort, Plan, Decision, Session, Artifact, Run (and optional Agent, Review) with refs and validation; MCP + generated TS + GraphQL over one model; append-oriented agent writes into the graph. Target: multi-week agentic work in repos. Strength: differentiated category, testable MVP, uses every Flatbread primitive; clarifies the GraphQL pivot as "one of several query adapters." Risk: needs a minimal credible write path and a schema flexible enough for real harness diversity without dissolving into bespoke configs.
Relational sketch (conceptual):
flowchart LR
Effort --> Plan
Effort --> Session
Plan --> Decision
Session --> Artifact
Session --> Run
Decision --> Artifact
Flatbread is the relational layer for agent efforts in git. Model efforts, plans, decisions, sessions, and artifacts as typed flat-file collections. Query them with MCP, GraphQL, or generated TypeScript. Catch broken references before the next session loses the thread.
Why not A: Avoids building a competing memory OS and UI; stays compositional with existing harnesses.
Why not only B: B is the implementation spine of C; C adds the Effort noun and integrity guarantees that make the positioning legible and defensible during a pivot away from GraphQL-only.
Illustrative only—not a shipping API. Shows how collections and refs express the graph (paths and collection names are placeholders):
import { defineConfig, transformerMarkdown, sourceFilesystem } from 'flatbread';
export default defineConfig({
source: sourceFilesystem(),
transformer: transformerMarkdown({ markdown: { gfm: true } }),
content: [
{
path: '.flatbread-proof/efforts',
collection: 'Effort',
refs: { owner_agent: 'Agent' },
},
{
path: '.flatbread-proof/plans',
collection: 'Plan',
refs: { effort: 'Effort' },
},
{
path: '.flatbread-proof/decisions',
collection: 'Decision',
refs: { effort: 'Effort', plan: 'Plan', session: 'Session' },
},
{
path: '.flatbread-proof/sessions',
collection: 'Session',
refs: { effort: 'Effort' },
},
{
path: '.flatbread-proof/artifacts',
collection: 'Artifact',
refs: {
effort: 'Effort',
session: 'Session',
source_decision: 'Decision',
},
},
{
path: '.flatbread-proof/runs',
collection: 'Run',
refs: { session: 'Session', effort: 'Effort' },
},
{
path: '.flatbread-proof/agents',
collection: 'Agent',
},
],
});Frontmatter fields (e.g. id, status, blocking, decided_at, tool) would drive filters; body markdown holds narrative. This sketch pressure-tests PMF priorities: stable id semantics, duplicate detection, missing-ref diagnostics, and typed config for collections.
- MCP server — tools: list collections, query with existing filter DSL, expand refs, fetch artifact body; primary agent-tailored surface for the GraphQL pivot.
- Generated TypeScript accessors — typed queries aligned with codegen (
packages/codegen). - GraphQL — keep for humans, Studio, and apps already on Apollo.
- Append / deposit API — narrow writes: create or append artifact rows with validation (no general transactional DB).
- Watch mode — content reload without full process restart (audit gap on local loop).
- Conventions preset — optional mapping from common paths (
AGENTS.md,SKILL.md,.cursor/rules,.handoff/*,.GCC/branches/*,.cs/*) into derived or linked collections without forcing migration day one.
- Query surfaces: GraphQL becomes explicitly one adapter over a single relational model; MCP and TS carry agent workflows.
- Writes: First-class but scoped—append-only artifact deposits + validation—not OLTP.
- ICP shift: From "relational markdown for sites" toward "multi-session agentic efforts in repos"—faster-moving buyer and clearer wedge than generic flat-file DB comparisons (per audit).
- Roadmap elevation: ID normalization, relation validation, and watch mode rise from hygiene to product blockers for this use case.
Cross-check against What Not To Build Yet and related warnings.
| Audit constraint | Relationship to Posture C (Proof) |
|---|---|
| Do not build hosted CMS, dashboard, or editing UI yet | Respects. MCP/CLI/codegen only; no mandated admin UI. |
| Do not compete with full databases on transactions, auth, permissions, high-scale writes | Respects if writes stay append-oriented, local, and validation-focused; stretches if users demand multi-user locking or roles—then explicitly out of scope. |
| Do not over-invest in many source plugins before local filesystem relational workflow is excellent | Respects if Proof ships as one filesystem preset + optional path mappings; stretches if many SaaS sources are added prematurely—avoid. |
| Do not keep GraphQL as the only story | Aligns. MCP + TS are first-class in this thesis. |
| Do not add complex migration systems before schemas, IDs, validation, exports, watch | Aligns. Proof assumes those foundations land first; reopens migration only as import from handoff/GCC folders once core is stable. |
| Avoid database replacement framing until write path exists | Stretches language—"relational layer" must stay precise: not a serverless Postgres; reopens the need for a documented, minimal write story before marketing breadth. |
Section 8 (schema sketch) and 9 (surfaces) should stay synchronized with audit honesty: no promise of generic mutations until shipped.
Before roadmap commitment:
- Preset wire-up — Point Flatbread at an existing agent folder in a real repo (e.g.
.agents/or documented handoff layout). Can one MCP query return all blocking decisions for the current effort with plan title without custom scripts? - Adversarial schema — One Proof schema against three layouts: Claude Code-oriented tree, Cursor rules + skills layout, GCC
.GCC/layout. Where does a single schema break? What mapping layer is minimal? - Token budget — Compare cold-start context stuffing vs Flatbread-mediated retrieval over a multi-session effort; benchmark against published handoff savings orders (e.g. AgentHandoff's reported range) as a directional bar, not a guarantee.
- Hosted editing UI or CMS.
- Mandatory vector index inside core (optional plugin later).
- Arbitrary update/delete mutation API on arbitrary fields.
- New source plugins for proprietary SaaS artifact stores before filesystem Proof is excellent.
- Replacing Claude Code, Cursor, or Codex harnesses—Flatbread should compose, not compete.
Thesis: Flatbread should own typed, validated, queryable effort graphs in git, exposed to agents primarily via MCP and generated TypeScript, with GraphQL as a parallel adapter. Next step: run the three validation experiments in §12; let results set the minimum mapping layer and write-scope for v1 of the preset.
- Wu et al., Git Context Controller (GCC).
- Vasilopoulos, Codified Context: Infrastructure for AI Agents in a Complex Codebase.
- Lore: Repurposing Git Commit Messages as a Structured Knowledge Protocol (arXiv).
- Claude Code Skills documentation.
- AgentHandoff, claude-sessions, agmem, Agentic Beacon.
- Vault / knowledge MCP examples: markdown-vault-mcp, engraph.