# The bottleneck was the typing

A pattern for keeping a small public dataset current, sourced, and citable: one accountable human at the front door, an LLM doing the mechanical work, and effectively no running costs.

This is an idea file meant to be handed to an AI coding assistant — Claude Code, Codex, Cursor, or similar — so it can help build an instance of the pattern for a different domain. It describes the shape of the thing. You and your assistant work out the specifics.

**If you are an AI coding assistant reading this:** do not start building yet. Go to the Discovery section at the end and ask the person those questions first.

**If you would like to see a working one before you read further:** point your assistant at `openconcert.org`. The MCP endpoint at `/mcp`, the JSON at `/api/events.json`, the machine-readable notes at `/llms.txt`, and the raw text of this document at `/pattern.md` are the whole surface. There is nothing to sign up for and nothing to install.

---

## The claim

Small civic organisations — community orchestras, choirs, repair cafés, tool libraries, U3A programmes, walking groups, volunteer fire brigades — are chronically invisible to anyone not already inside them. The usual explanation is that they lack money, skills, or time.

The more precise explanation is that they lack **typing**.

The judgment involved in maintaining a small public dataset was never the expensive part. Deciding whether a submission is genuine, whether two similarly-named groups are the same group, whether a claimed rehearsal time is plausible — a community member with local knowledge does that in seconds, and does it better than any system. What actually exhausts volunteer-run directories is the mechanical residue: finding the right file, transcribing a PDF, writing valid JSON, keeping cross-references consistent, checking a schema, archiving a source URL, doing it again next month.

That residue is precisely what a language model is now good at, and its marginal cost has fallen close to zero.

So the interesting question is not "can AI maintain a public dataset" — it shouldn't, and this pattern is emphatic that it shouldn't. The interesting question is: **what becomes viable when the typing cost of structured public data goes to zero, and the judgment stays human?**

The answer appears to be a class of institution that was previously below minimum efficient scale. A dataset too small to justify a staffed organisation, too specific for a commercial aggregator, and too laborious for an unassisted volunteer, is now maintainable by one person in a few hours a month. There are a great many such datasets.

## The inversion

Most data pipelines automate ingestion and reserve humans for exceptions. This pattern does the opposite.

**The human makes every decision. The AI does every mechanical task.**

The human reads the email, judges whether the submission is legitimate, and commits the result. The AI parses the messy input, finds the right file, writes valid JSON, resolves cross-references, validates against the schema, archives the source to the Internet Archive, and presents a diff.

The maintainer is a community member, not a service operator. Their inbox is the project's front door. The role is closer to a parish secretary or a library volunteer — someone who knows the local groups and helps newcomers in — than to a content moderator reviewing automated output.

**Human at the gate, not human in the loop.** The AI works in the back office; the person stands at the door.

The distinction matters most for data about people and communities who did not ask to be indexed. A choir secretary emailing a concert date is writing to a person, not submitting to a pipeline. That relationship — a known, accountable member of the community the data is about — is what makes the data worth citing at all.

## The problem it addresses

Information about small volunteer organisations lives in places that require already being inside: members-only mailing lists, Facebook events behind a login, PDF programmes, paper flyers, small-CMS websites that block crawlers, committee secretaries' inboxes. None of it is secret. Much of it is private by inaccessibility.

When an outsider asks a search engine or an AI assistant, that gap gets papered over with confident guesses. Two ensembles are merged into one. A founding date is inverted. A concert is attributed to the wrong group. The model is not lying; it is interpolating across absent data. Scrape-and-publish aggregators do not help, because a machine's interpretation of another machine's scrape carries no accountability and no one to write to when it is wrong.

The same gap erases the past. Volunteer organisations produce paper records — programmes, member lists, repertoire histories — that vanish into filing cabinets when their keeper moves house. Decades of local cultural history are lost quietly, generation by generation, because nobody built a place to put them.

### What scrapers cannot see

This is not an anti-scraping position. Where a site permits crawling, the maintainer fetches the page, archives it, and writes a record like any other. `robots.txt` is respected; if a site says no, that is the answer.

What is striking is how much civic data sits outside what any scraper could reach *even with full permission*.

- **Mailing lists.** Much of community life runs on them. "Open rehearsal on Sunday." "We are looking for a new alto." These go only to subscribers. **You cannot join what you cannot see** — and what gets posted on those lists is usually the way in: the come-and-try day, the open rehearsal, the call for a new alto. The would-be alto never sees the call for the alto. That double bind is structural, and it is one of the quieter reasons that welcoming organisations stay small.
- **Closed platforms.** Facebook events, members-only groups, WhatsApp threads. Online in some sense; undiscoverable to anyone not already on the platform following the right account.
- **Bot-shielded small sites.** Wix, Squarespace, and Cloudflare-fronted club sites increasingly block crawlers by default. The organisation would happily be indexed; its hosting says no on its behalf.
- **PDFs, paper flyers, photographs of noticeboards, forwarded screenshots.** Offline, or close enough.

A maintainer's inbox sits *across* that boundary. A person can subscribe to a community mailing list as a normal member. A secretary can email a PDF. A friend can paste a Facebook event into a chat. None of this circumvents anyone's access controls: the maintainer is invited in, and re-publishes, with attribution, the public-facing parts of what the community already chose to share.

Without that bridge, the most accessible and most welcoming organisations in a community remain the least findable.

## Voluntary legibility

There is an old and well-founded suspicion of making communities legible to outside systems, because historically legibility was imposed from above and served the party doing the imposing.

This pattern runs that in reverse. Nothing is compelled and nothing is extracted. An organisation chooses to be findable, sends what it already has, and can correct or contest any record by writing to a person who will answer. The legibility is opt-in, bottom-up, attributed, and revocable in practice — the maintainer will remove a record if the organisation asks. What is published is what the organisation already puts on its own noticeboard, moved somewhere a stranger can find it.

The failure mode of imposed legibility is that the map flattens what it cannot see and then the map is enforced. The corrective here is the same one that makes the data citable: a named human is accountable for each record, and the source of every claim travels with it.

## No accounts, no forms, no friction

There is no submitter login, no registration, no portal. The maintainer is the only authenticated user of the system and the only one who needs to be.

Submitters send whatever they already have. They forward the email they already wrote to their own members. They send the PDF they already gave to the printer. They photograph the flyer already pinned to the church-hall noticeboard. The ingestion skill accepts whatever shape the source arrives in; the maintainer decides whether to include it.

Nothing is treated as proof of identity or authority — not a verified email address, not a claimed role. The provenance fields record the channel and timestamp; the named maintainer's judgment is what stands behind the record.

Adding a registration step would move work onto the people the pattern exists to serve — committee volunteers with no time for it — in exchange for a weaker signal than the one already present. Friction on the submitter side is a cost, not a feature.

## Not a destination

A conventional directory is built for human visitors who browse, bookmark, and search it. Traffic is the metric. You invest in UI, SEO, engagement.

Civic data infrastructure is built for machines to read so that humans can find things. The end audience is still human — someone looking for a choir to join, a researcher dating a 1987 concert, a journalist checking who conducted what. They simply do not arrive by typing the URL. They arrive via an AI assistant, a search result, a calendar subscription, or an app that pulled the JSON.

Nobody will bookmark the site. AI assistants will read it. Developers will embed it. Crawlers will ingest the structured markup. The humans show up at the other end.

This reframes every product decision. The MCP endpoint matters more than the homepage. A `verified: true` field with a dated archive URL is worth more than any amount of visual polish.

## There is almost no code

This is the part that surprises people, so it is worth being precise rather than rhetorical.

The reference implementation deploys a single stateless worker that imports **exactly one third-party library** — an HTTP router. There is no database, no ORM, no migrations, no authentication, no sessions, no queue, no client-side build step, and no server to keep alive. Nothing in the system holds state between requests. A build step folds every JSON file into one module; the running service reads that module and formats it.

What the repository actually contains, in rough order of importance:

| Layer | What it is | Size |
|---|---|---|
| The data | One JSON file per record, in git | ~935 files |
| The schema | JSON Schema, enforced as a hard commit gate | 553 lines of JSON |
| The skill layer | **Instructions written in English** for an LLM to follow | 1,989 lines of prose |
| The output layer | Functions that turn the bundle into HTML, JSON, iCal, RSS, JSON-RPC | ~8,700 lines of TypeScript |
| The scripts | Validate, bundle, archive, geocode, smoke-test | ~3,200 lines |

Two things in that table are worth dwelling on.

**The layer that does the actual work is prose.** Thirteen instruction files — ingestion, inbox triage, routine maintenance, pre-flight checks — written in English, describing how to classify an incoming email, where to find the right file, how to extract a movement marking from a work title, when to refuse an edit and draft a reply to the sender instead. It is not code. It is the sort of document you would write for a new volunteer, and it happens to also be executable now. Anyone who can write a clear procedure can maintain it. Nobody has to be able to program to change how the system behaves.

**The output layer looks large but is not deep.** Most of those TypeScript lines are HTML templates and per-surface formatting — the same records rendered five ways. There is no business logic in there to get wrong, because there is no business. The most algorithmically substantial file in the entire project is a QR-code encoder — Galois-field arithmetic, Reed–Solomon error correction and all — written from scratch so that a printed concert programme can carry a scannable link. The most code-like code in the project exists to reach paper.

The implication for anyone building one: **the hard part is not the software.** It is deciding what a record is, what the repeated units are, what you can resolve identities against, and where the sensitive-data boundary sits. Get those right and the rest is formatting. Get them wrong and no amount of engineering rescues it.

This is also why "please open-source the code" is the wrong ask, and why the code being closed costs a would-be builder nothing. There is very little there that would transfer. The transferable parts are the schema shape, the invariants, and the skill structure — all of which are in this document, below.

## The two promises

Everything else is negotiable. These are not.

**1. Every record cites its source.** No ghost data. A record without provenance does not ship.

**2. Every cited URL is snapshotted.** The source page is pushed to the Internet Archive and the snapshot URL is stored on the record itself. Small-organisation websites go dark constantly; when one does, the citation still resolves.

Provenance fields encode the chain so any reader can follow it:

- `imported_from` — where the information came from
- `last_checked` — when it was last verified (ISO date)
- `checked_by` — role, channel, and timestamp (no personal names — the channel and timestamp are the audit trail)
- `verified` — did a responsible party at the organisation confirm this record?
- `archive_url` + `archived_at` — the snapshot, so the evidence survives the source

A `verified: true` record with a dated archive URL is a materially stronger citation than an unverified scrape, and the difference is legible to a machine.

**PII discipline is not negotiable either.** The data is openly licensed, which means anything written to a file is published. Record the email's date and time as the audit trail — enough for the maintainer to find the original in their own inbox. Do not record sender names, addresses, or claimed roles unless that information is already on the subject's own public website. The `verified` flag carries the signal; who confirmed it is private correspondence.

## What it costs

| | |
|---|---|
| Hosting | $0 — a stateless worker on a free tier, 100,000 requests/day |
| Source control and CI | $0 |
| Inbound email | $0 — any address that routes to an ordinary mailbox |
| Domain name | ~$15/year, and optional |
| LLM usage | A consumer subscription, or a few dollars a month in API credits |
| Maintainer time | A few hours a month |

The infrastructure genuinely is free, and that is a design constraint rather than a boast: **free infrastructure for organisations that are themselves free to join.** A volunteer-run choir should not face a hosting bill on top of everything else it is already doing.

If traffic ever exceeds the free tier, the architecture is already stateless. Moving to a paid tier or a different host is a configuration change, not a rewrite.

## What it looks like after a while

The reference implementation, as at August 2026, after about three months of part-time maintenance by one person. Current figures are always live at `/api/events.json` and `/stats`.

- **11 ensembles**, **407 events**, **458 composers**, **66 venues**
- Events span **1996 to 2026** — 31 distinct years, of which **380 events are already past**. The archive is the majority of the dataset, and grows more useful the longer it runs.
- **326 events carry a full transcribed programme**: 1,697 individual work performances, covering 1,221 distinct works. Most of that came off paper.
- **407 of 407 events cite a source.** Not a target — a schema gate.
- Every event record whose source is a URL has an Internet Archive snapshot stored on it. Same for every ensemble profile. Composer records are at 199 of 205, with the remainder in the periodic sweep.

The scale is worth noticing in both directions. Four hundred records is trivially small by any commercial standard, and it is *also* three decades of local repertoire history that existed nowhere else in structured form and was, in the ordinary course of events, going to be thrown out. It took one person a few hours a month for a season.

There is a public telemetry page at `/stats` reporting which AI crawlers actually arrive, which machine-readable tools actually get called, and where referrals come from. It is worth publishing that honestly, including when the answer is unflattering, because it is the only evidence anyone has about whether this kind of thing works. It should also be read carefully: **being crawled is not being cited.** The first is measurable and the second, so far, largely is not.

## Architecture

Three layers.

**The substrate.** A directory of JSON files in a git repository — one file per entity. Files are the source of truth; there is no database. Entities cross-reference each other by slug rather than by embedded copy, and where possible carry canonical external IDs (an open composer registry, OpenStreetMap for venues, Wikidata for anything well-known) so that the same thing in two places lines up automatically and two near-identical things stay distinct. Schema validation is a hard commit gate: what fails does not ship.

**The ingestion layer.** A set of LLM skills — slash commands, agent instruction files, whatever your tooling calls them — that turn messy input into clean records. An email arrives, the maintainer forwards or pastes it, and the skill classifies the input, dispatches to a new-entity or update path, finds the target file, applies the change, records provenance, validates, archives cited URLs, and shows a diff. The maintainer reviews and commits.

Critically: **when something does not smell right — an authority claim that does not match the source, a high-stakes change with no documentary evidence, a contradiction with the existing record — the skill drafts a reply to the sender instead of applying the edit, and the record stands unchanged.** Smoothing past ambiguity is the failure mode the whole pattern is built to avoid.

**The output layer.** A build step bundles every JSON file into a single module the worker imports at deploy time. From that one bundle the worker serves HTML with structured markup embedded, a REST JSON API, an MCP endpoint, calendar and feed subscriptions at several scopes, short links, QR renders for printed inserts, and the crawler files. Adding a surface means writing a new renderer over the same bundle. It is never a new pipeline.

### Which surface is for whom

- **HTML with embedded structured data** is for people arriving via a search engine or an LLM browsing the web. The markup is what lets an intermediary ingest the page rather than guess from prose.
- **REST JSON** is for developers and embeds. Single-entity responses should carry enough context — the event *and* its ensemble — that a simple caller never needs a second request.
- **MCP** is for AI assistants used as tools. This is the strongest form of citable data: the assistant is not guessing from a scrape, it is querying structured records with the provenance fields sitting right alongside the content.
- **Calendar and feed subscriptions** are for humans who want the data to come to them.

## Operations

**Inbox.** The primary channel is email, and submissions come from three kinds of sender: organisations writing in directly, readers correcting records, and community mailing lists the maintainer belongs to as an ordinary member. The last matters more than it sounds. The maintainer checks on a regular cadence, filters noise, and walks through genuine submissions one at a time. A skill handles filtering, fetches thread content, classifies each submission, and dispatches. Actioned threads are tracked by thread ID in a committed state file — cross-device, no personal data.

**Ingest.** For a new entity or a batch, the maintainer pastes source material and the skill parses it into JSON, resolves cross-references, validates, archives, and writes files. The maintainer reviews the diff.

**Edit.** For updates, the skill finds the target file, applies the change, updates provenance, re-validates, and re-archives if the source changed. One email, one targeted edit, one commit.

**Publish.** A pre-flight script validates everything, regenerates the bundle, typechecks, and smoke-tests the output surfaces. If it passes, deploy — a single command.

Two things here are non-negotiable: **the maintainer reads each email, and the maintainer writes the reply.** The AI summarises threads, drafts, finds files, formats data. But the message that goes back to a choir secretary is from a person, in a person's voice, to a person. An inbox that auto-responds would destroy the one property that makes any of this worth citing: that someone was paying attention.

## The relationship dimension

An unanticipated consequence of having a person at the gate: the data relationship becomes a community relationship.

A secretary emails a correction; a person reads it and applies it; that person replies to confirm. The secretary learns there is a real human maintaining this. They start to think of the maintainer as local infrastructure — someone who, like the council's parks team, keeps a civic thing working. They mention it to colleagues. They send updates before the next season without being asked. They introduce the maintainer to other organisations.

None of this happens when a form submission enters a pipeline. It happens because the relationship is between people.

The inbox is not a bottleneck. It is the point of contact with the community the data is about.

## Scope: one domain, not one city

This is the design decision most likely to be got wrong, and it is worth stating flatly.

**A dataset should be global in its domain and narrow in its subject.** One registry covering community orchestras everywhere is far more useful than forty city registries covering community orchestras separately — because fragmentation destroys exactly the property the pattern exists to create. An assistant that must guess which of forty endpoints to consult will consult none of them. A schema that is multi-tenant from the first day costs nothing extra and prevents the split later.

So if the domain already has an instance and you want to cover a new city, the useful move is almost always to add the city to the existing registry and become the person who maintains it, not to stand up a parallel one. Coverage is a maintainer problem, not an infrastructure problem.

**A new instance earns its existence by covering a different subject**, where the repeated units, the schema, and the sources are genuinely different: repair cafés, community gardens, tool libraries, amateur sports leagues, historical societies, volunteer rescue units, men's sheds, adult-education programmes. Each of these has the same shape — invisible to outsiders, distributed across inboxes and closed platforms, and possessed of a past that is being thrown away — and each needs its own schema, because a repair café is not a concert.

## Iteration

Start with the minimum data layer that achieves the outcome, and no more. Get the data right before reaching for a frontend. Because the audience arrives via assistants and search engines, plain server-rendered text with structured metadata matters far more than visual design. A new entity, a new field, a new surface earns its place when a real submission or a real reader makes one obvious — not before. If an addition would take a weekend, do it; if it needs a roadmap, the substrate is probably not the right place for it.

The skills should learn as they go. When the same edge case has come up three times, or the maintainer keeps applying the same tweak after a skill finishes, the AI should propose a small plain-language amendment to the skill's own instructions for the maintainer to approve. The instructions get sharper the longer they run, and nobody has to configure anything by hand. This is only possible because the ingestion layer is prose.

The two promises stay fixed. Everything else is allowed to move, slowly, in response to use.

## What this does not solve

**It is not a comprehensive directory and should not try to be.** One maintainer with an LLM can keep a few hundred to a few thousand records honest. Past that, the human-judgment-per-record property breaks and you are back to the scrape problem the pattern was designed to avoid. Attempting completeness produces exactly the large, stale, uncorrectable directory this exists to replace. **The ceiling is a feature**, and a real limit: this is a pattern for many small datasets, not one big one.

**It does not find organisations the maintainer is not already adjacent to.** If nobody has mentioned a group, it will not be listed. The pattern lowers the barrier for newcomers to *find* organisations; it does not discover them.

**It does not survive the maintainer disappearing.** Trust attaches to a named person. If they go quiet and nobody picks up the inbox, the project goes quiet. Succession is a live operational concern, not a future one.

**It does not fix bad source data.** If a secretary writes 8pm when the concert is at 7pm, the record says 8pm until somebody notices. Provenance tells you where a claim came from. It does not tell you the claim is true.

**It does not compel anyone to cite the data.** Publishing structured markup, an MCP endpoint, and a sitemap makes citation easier and more likely. It guarantees nothing. Some assistants will keep preferring their own scrapes, and there is currently no reliable way to observe whether the data influenced an answer.

**And it does not answer who ought to pay for this.** Verified, attributed, human-checked ground truth is a public good that AI systems increasingly depend on and nobody is funding. This pattern demonstrates that a volunteer can supply it for a domain at close to zero cost, which is an existence proof, not a theory of provision. Whether unpaid supply is enough, at what scale it stops being enough, and what should happen then, are open questions. Volunteer effort has historically underwritten a great deal of infrastructure of exactly this kind; that is an encouraging precedent and a fragile one.

**It does not replace showing up.** Knowing an ensemble exists is not walking into a rehearsal room. The job is to get a would-be member to the door, not through it.

---

## Build notes

Enough to start. Not a framework — a shape.

### Invariants, as assertions

Enforce these in the validator, not in prose:

1. Every record has a non-empty source field. Fail the build otherwise.
2. Every slug matches `^[a-z0-9][a-z0-9-]{1,63}$` and the filename agrees with the record's own ID.
3. Every cross-reference resolves to a file that exists.
4. Every URL written to a source field acquires an archive snapshot, or lands in a backlog that a periodic job clears.
5. No field contains a personal name or email address unless that information is already on the subject's own public website.

### A record, in outline

The domain here is concerts; substitute your own. What matters is the shape: a stable ID, a status lifecycle, cross-references by slug, and a provenance block that is never optional.

```json
{
  "id": "northbrook-community-orchestra/2026-09-14-spring-concert",
  "ensemble": "northbrook-community-orchestra",
  "status": "programme-published",
  "date": "2026-09-14",
  "start_time": "19:30",
  "timezone": "Australia/Sydney",
  "venue": { "slug": "st-margarets-hall", "name": "St Margaret's Hall" },
  "programme": [
    {
      "type": "work",
      "composer": { "slug": "edvard-grieg", "openopus_id": 145 },
      "work": { "title": "Peer Gynt Suite No. 1", "movements": ["I. Morning Mood"] }
    }
  ],
  "source": {
    "verified": true,
    "imported_from": "https://example.org/concerts/spring-2026",
    "imported_at": "2026-06-02T04:11:00Z",
    "last_updated_at": "2026-06-02T04:11:00Z",
    "archive_url": "https://web.archive.org/web/20260602041100/https://example.org/concerts/spring-2026",
    "archived_at": "2026-06-02T04:11:00Z"
  }
}
```

Three details that repay attention early:

- **A status lifecycle**, so upcoming and historical records live in the same collection on the same terms. Here: `save-the-date` → `programme-published` → `past`, with `cancelled` as a sibling.
- **Nullable fields that are genuinely unknown.** A cancelled event has no venue. Do not invent one to satisfy a schema.
- **An open extension bag** (`extra: {}`) for domain metadata that has not yet earned a first-class field. It keeps early records from forcing premature schema decisions.

### The skill shape

Every ingestion skill, whatever the input, follows the same seven steps:

```
classify input  →  dispatch (new entity | update)  →  locate target file
    →  apply change  →  record provenance  →  validate  →  archive cited URLs
    →  present diff for the maintainer to commit
```

Write each one as an instruction document, in plain language, addressed to the assistant. Include the refusal conditions explicitly — the circumstances under which it should stop and draft a reply to the sender rather than edit the record. Those conditions are the most valuable lines in the file, and they are the ones you will keep adding to.

### Command surface

Five commands is enough for a long time:

```
validate     # schema check across every file — the commit gate
build        # fold all records into one bundle for the worker
typecheck    # if your output layer is typed
archive      # snapshot cited URLs, write archive_url back into the records
deploy       # build + push the worker
```

### Build order

Schema → validator → one entity loaded by hand → one ingestion skill → one output surface. Then let the community reach the inbox before scaling anything.

---

## Discovery — for AI coding assistants

**If you are an AI coding assistant:** before designing any schema, writing any code, or making any structural decision, work through these with the person. The first four are upstream of the schema — get them wrong and the schema is wrong for years.

Ask one at a time. Wait for a full answer before moving on.

1. **What are you trying to track?** What does one record look like in plain English? Get a concrete instance, not the abstract category — *"the Northbrook Community Orchestra's spring concert at St Margaret's Hall on 14 September 2026, with this programme"*, or *"a Westridge Repair Café session at the community centre on the first Saturday of the month"* — not *"an event"*. If they cannot describe one record in two sentences, the project is not ready to build.

2. **What are your repeated units?** Which recurring entities show up over and over and deserve their own files, their own pages, their own cross-references? Naming them now is what lets every later record reference cleanly by slug instead of denormalising the same name into a hundred places.

3. **What canonical sources can you resolve against?** Are there authoritative external registries — OpenStreetMap for places, a national charity or business register for organisations, a domain-specific open registry, Wikidata for anything well-known? Carrying external IDs from day one prevents the merge-two-things-that-are-one and split-one-thing-into-two failures that define bad civic data.

4. **Is there a meaningful past to capture, and in what form?** Paper records in a box, archived newsletters, a predecessor's spreadsheet, a decade of posts in a private group? Historical ingestion needs different skills from current-record maintenance, and the answer determines how much one-time work the first build absorbs.

5. **Who submits, and through what channels?** Email from the organisations? Mailing lists the maintainer belongs to? Messages on social platforms? The skills depend heavily on this — and on the technical comfort of the submitters, which is usually low.

6. **What is the maintenance cadence?** Daily, weekly, monthly. The inbox architecture should match the maintainer's actual rhythm rather than assume one, and the promise made to the community should match it too.

7. **Where is the sensitive-data boundary?** What in this domain is personal, organisationally sensitive, or otherwise not for public commit? Open licensing means anything in a file is published. Establish this before the schema exists and hard-code it into every skill, so the project cannot drift past it.

8. **What does the first version look like, and who is the first real user?** Not the long-term vision. Name an actual person, say what they will do with the first deployment, and identify how they will reach it. The schema, the surfaces, and the build order should serve that person before anyone else.

Only once these have clear answers: propose a schema, a file layout, the first one or two skills, and a phased plan. Show all of it to the person before writing files. Adjust. Then build in the order given above.

The pattern is reusable. The instance is theirs to design.

---

*OpenConcert — a registry of live community music, and the reference implementation of this pattern — is at [openconcert.org](https://openconcert.org). Its data is published under CC BY 4.0. The plain-text version of this document lives at [openconcert.org/pattern.md](https://openconcert.org/pattern.md).*
