Kayak: one contract, every face
I’ve encountered this same bug more than once, in more than one language. Someone adds a listing sorted by some field, everything passes, and the EndpointOne specific web address a service answers requests on, such as the one that lists files. goes live. Then, Months later the table holds a few million rows and that endpoint is the slowest route in the service, because someone had forgotten to add an index. The database did what was asked, but it still had to read every row.
The other bug I’ve seen happens just as often but is less obvious. An API (application programming interface)The set of requests one program accepts from another. A web API is how apps, scripts and AI agents ask a service to read or change its data. has more than one face: RESTThe most common style of web API. Each kind of thing gets a web address, and you read or change it with standard HTTP requests. routes, GraphQLA query language for APIs in which the caller asks for exactly the fields it wants and gets them back in one response. SchemaThe declared shape of a database or an API: which tables or types exist, what fields they have, and how they can be looked up., OpenAPIA standard format for describing a REST API. Tools read an OpenAPI document to produce documentation, test tools and client code. schemas, or an MCP (Model Context Protocol)The standard way AI coding agents such as Claude Code connect to outside tools and data. server all hand-rolled but serve the same route but only 1 of them has a different header. Oops!
Kayak is the tool I wrote to stop both the bellyaching over schemas, and to prevent these silent mistakes from creeping back into the codebase. The idea for this is simple:
Declare the API once, check the declaration against the real indexes, and generate every surface from that singular object. Surfaces that come from the same object cannot disagree.
The name comes from the problem it addresses. Kayak exists to stop Schema driftWhen the database, the code and the documents that are supposed to describe the same data slowly stop agreeing with each other..
What Kayak is
Kayak is a library written in RustA programming language known for being fast and for catching whole classes of bugs before a program ever runs. beside a small CLI (command-line interface)A program you use by typing commands into a terminal rather than clicking through windows. for APIs backed by SurrealDB. You describe what an API exposes in a Contract (Kayak)In Kayak, the single checked-in declaration of everything an API offers, from which the documents, the clients and the live server are generated.: which tables, which columns, what callers may filter and sort on, and which actions exist beyond list and get. The contract never restates the schema. It points at tables and columns by name, and Kayak resolves those names against the schema you already define in code with surql-rs, the client library I wrote about back in April.
From one object Kayak produces an OpenAPI 3.1 document, a
GraphQL SDL (Schema Definition Language)The plain-text format that describes a GraphQL API: its types, their fields, and the queries and changes it accepts., an MCP tool manifest, and clients in Rust,
TypeScriptJavaScript with type annotations added, widely used for web apps and servers. The types catch many mistakes before the code runs., PythonA popular, easy-to-read programming language used for everything from small scripts to data science. and Go. The clients are light so consumers do not have to wrangle code.
The TypeScript client uses the built-in fetch API, the Python
client uses std library only, the Go client is written against net/http,
and the Rust client leverages reqwest and serde libs.
With the runtime feature enabled, Kayak also serves the contract.
When hosted, Kayak provides a live GraphQL schema, a REST router, an operator console, and a dispatcher that clients can call, and all of
them enforce the same rules because all of them go through Kayaks Dispatcher (Kayak)The part of Kayak's runtime that checks each request against the contract, applies limits and permissions, and only then calls your code..
Defining a contract
Say you run a small helpdesk. Tickets live in a SurrealDBAn open-source database that stores ordinary tables, linked records and nested documents in one engine. Kayak, Copal and Antumbra are all built on it. table, and the schema is already code:
use surql::schema::{
datetime_field, index, int_field, string_field, table_schema, FieldBuilder, TableMode,
};
let built = |b: FieldBuilder| b.build_unchecked().unwrap();
let ticket = table_schema("ticket")
.with_mode(TableMode::Schemafull)
.with_fields([
built(string_field("tenant_id")),
built(string_field("title")),
built(string_field("status")),
built(int_field("priority")),
built(string_field("assignee").nullable(true)),
built(string_field("internal_note").nullable(true)),
built(datetime_field("created_at")),
])
.with_indexes([
index("idx_queue", ["tenant_id", "status", "created_at"]),
index("idx_assignee", ["tenant_id", "assignee"]),
]);The contract says what the API does with that table. Contracts are plain data, so they can live in Rust next to the service or in a JSONA plain-text format for structured data, built from named fields and lists, that almost every programming language can read and write. file next to the build. This is the JSON form:
{
"name": "helpdesk",
"version": "0.1.0",
"auth": { "kind": "bearer" },
"resources": [
{
"name": "tickets",
"table": "ticket",
"fields": [
{ "column": "title" },
{ "column": "status" },
{ "column": "priority" },
{ "column": "assignee" },
{ "column": "created_at" }
],
"pinned": ["tenant_id"],
"filterable": ["status", "assignee"],
"sortable": ["created_at"],
"max_page_size": 50,
"actions": [
{
"name": "close",
"method": "POST",
"path": "/{id}/close",
"input": [{ "name": "reason", "kind": "string" }],
"output": "none"
}
]
}
]
}Read it top to bottom and it's the whole promise the API
makes. Every read is pinned to the caller's tenant_id, which
the server binds and a caller never sends. Callers may filter
on status and assignee, sort on created_at, and page at
most 50 rows at a time. There is one action beyond list and
get, which closes a ticket. And internal_note isn't listed,
so it doesn't appear on any generated surface. Nothing is
exposed until it has been named.
Generating is one command. The schema file is the same table definitions serialized to JSON by whichever service owns them:
kayak generate --contract contract.json --schema schema.json --out generatedwrote generated/client.go
wrote generated/client.py
wrote generated/client.rs
wrote generated/client.ts
wrote generated/mcp-tools.json
wrote generated/openapi.json
wrote generated/schema.graphqlThis is the GraphQL half of what came out. The page size, the filters and the sort enum all trace straight back to lines in the contract:
type Ticket {
id: ID!
title: String!
status: String!
priority: Int!
assignee: String
created_at: DateTime!
}
enum TicketSort {
CREATED_AT_ASC
CREATED_AT_DESC
}
type Query {
tickets(limit: Int = 50, cursor: String, status: String, assignee: String, sort: TicketSort): TicketPage!
ticket(id: ID!): Ticket
}
type Mutation {
ticketClose(id: ID!, reason: String): Boolean!
}And this is the Python client in use. The constructor takes a token because the contract says callers authenticate with a Bearer tokenA secret string a caller sends with every request to prove who it is. Whoever holds it is let in, so it has to be kept private. as we defined it earlier:
from client import Client
helpdesk = Client("https://helpdesk.internal", token)
page = helpdesk.list_tickets(limit=20)
for ticket in page.items:
print(ticket.title, ticket.status)
helpdesk.close_ticket(page.items[0].id, {"reason": "duplicate"})The gate names the column
Now, how it catches those pesky bug I mentioned. Someone wants the queue sorted by priority, so they add it to the contract:
"sortable": ["created_at", "priority"]contract failed validation:
- resource tickets: sortable column priority is not reachable as an index sort suffix on ticket; some index must hold it with every earlier column pinned or filterable, or ORDER BY falls off the indexThat message is most of the reason Kayak exists. The rule behind it is the one the database
follows. An index can hand rows back in order only when every
column ahead of the sort column is held to a single value.
idx_queue is (tenant_id, status, created_at). The server
always pins tenant_id, and status is something a caller
can filter on, so a caller who filters by status gets tickets
back already in created_at order, straight off the index.
Nothing holds priority in a position like that. Sorting on
it means reading every ticket the TenantOne customer, team or workspace whose data is kept separate from everyone else's inside a shared service. has and sorting them
in memory, and with forty rows in a dev database you'd never
notice.
The fix belongs in the schema, not the contract. Add an index
that puts priority where the engine can walk it and regenerate it as before:
index("idx_triage", ["tenant_id", "status", "priority"]),And now the contract generates cleanly
The gate also covers the other claims a contract makes. A filterable column has to appear in some index. The pinned columns, as a set, have to lead an index, or every plain listing scans the table. Search gets checked too. A query can declare that it does Lexical searchSearch by the exact words in the text, as opposed to search by meaning. or vector search and name the index that answers it, and the gate checks that the index exists, holds the column, and is the right kind: Full-text indexA database index built for keyword search over text, so the documents containing given words can be found without reading them all. for lexical, a Vector indexAn index over embeddings that finds the stored items closest in meaning to a query, instead of matching exact words. such as HNSWHierarchical navigable small world, a popular kind of vector index that finds close matches quickly by hopping through layered graphs of neighbors. or DISKANNA kind of vector index designed to keep most of its data on disk rather than in memory, so it can hold very large collections. for vector. I added that rule after finding search paths in two of my own services that had run for months over columns no index could answer a Nearest-neighbour queryA search for the stored items closest to a given one, usually closest in meaning as measured between embeddings. on.
A static check can only reason about the definitions. When I
want the database's own opinion, kayak verify asks the Query plannerThe part of a database that decides how to answer a query, including which index to use or whether it has to read the whole table.:
kayak verify --contract contract.json --db ws://localhost:8000 --namespace app --database appIt runs an EXPLAIN for every filter, sort and search claim
in the contract and fails on any the planner answers with a
Table scanA database reading every row of a table to answer a query because no index lets it jump to the right rows. Fast on forty rows, slow on millions.. When everything holds, it prints
every filter, sort, and backing claim plans on its index.
The CLI needs the verify feature for this, since it's the
only part of Kayak that ever opens a database connection.
Breaking changes become an exit code
A contract is data, so two versions of it can be compared as
data instead of as text. kayak diff does that, and it exits
non-zero on anything that would break a client. This is a real
run where I renamed created_at to opened_at on the wire,
dropped the assignee filter, and added a reopen action:
kayak diff contract.json contract-v2.jsonBREAKING tickets: field created_at removed
compatible tickets: field opened_at added
BREAKING tickets: filter assignee removed
compatible tickets: action reopen addedA rename shows up as a removal plus an addition, which is what it is to a client that still asks for the old name. A text diff of two OpenAPI documents would show the same change as a wall of JSON and leave the judgment to whoever is reviewing it. Here the judgment is the Exit codeThe number a program hands back when it finishes. By convention 0 means success and anything else means failure., so CI (continuous integration)Automated checks that build and test every proposed change before it is merged, so problems are caught before they ship. can make it.
The rest of drift protection is important albeit a little boring. The generated files are checked in, and a test regenerates them and compares the bytes. This is the entire contract test in my penpal plugin, one of the services built on Kayak for Penpot:
use penpal::contract::{contract, schema};
#[test]
fn the_contract_validates_against_the_store_schema() {
let violations = kayak::validate(&contract(), &schema());
assert!(
violations.is_empty(),
"the contract drifted from the schema:\n{violations:#?}"
);
}
#[test]
fn the_artifacts_match_the_blessed_ones() {
let artifacts = kayak::generate_all(&contract(), &schema(), kayak::generate::TARGETS)
.expect("a validating contract generates");
let root = std::path::Path::new(env!("CARGO_MANIFEST_DIR")).join("generated");
std::fs::create_dir_all(&root).expect("the generated dir exists");
for (filename, content) in &artifacts {
let path = root.join(filename);
if std::env::var("PENPAL_BLESS").is_ok() {
std::fs::write(&path, content).expect("blessing writes");
}
let blessed = std::fs::read_to_string(&path).unwrap_or_else(|_| {
panic!("{filename} has no blessed copy; PENPAL_BLESS=1 cargo test")
});
assert_eq!(
content.trim(),
blessed.trim(),
"{filename} drifted from its golden; PENPAL_BLESS=1 if the change is intended"
);
}
}Re-blessing is a deliberate step, so a generated file only changes when someone meant it to, and the change shows up for review.
Serving the contract
The runtime feature is for when you want the
contract to do the enforcing as well as the describing.
Kayak still never talks to your database. You register ResolverThe function a service supplies to actually fetch or change the data for one operation. Kayak calls it only after the request has passed every check., which are async closures over your own data access, and Kayak calls them once a request has passed everything the contract says:
use std::sync::Arc;
use kayak::runtime::{Dispatcher, KayakContext, KayakError, ListOutput, Resolvers, RestRouter};
let resolvers = Resolvers::new()
.list("tickets", |ctx: KayakContext, args| async move {
// By the time this runs, args.limit is clamped to the page
// ceiling and args.filters and args.sort are allowlisted.
let tenant = &ctx.get::<Tenant>().unwrap().0;
let items = store::list_tickets(tenant, &args)
.await
.map_err(|e| KayakError::Internal(e.to_string()))?;
Ok(ListOutput { items, next_cursor: None })
})
.get("tickets", |ctx, args| async move { /* one ticket, or Ok(None) for a 404 */ })
.action("tickets", "close", |ctx, args| async move { /* close it */ });
let dispatcher = Dispatcher::new(Arc::new(contract), resolvers, vec![Arc::new(RequireTenant)])?;
let rest = RestRouter::new(Arc::new(dispatcher));Tenant is whatever type your HTTPThe protocol web browsers, apps and servers use to request and send data over the web. layer puts in the context,
and RequireTenant is MiddlewareCode that wraps every request on its way in and out, the usual place for checking who is calling, which tenant they belong to, and logging.. Middleware wraps every
operation on every face, which makes it the place for tenancy,
identity and logging:
use kayak::runtime::{BoxFuture, KayakError, Middleware, Next, Operation, Outcome, Payload};
struct RequireTenant;
impl Middleware for RequireTenant {
fn handle<'a>(
&'a self,
operation: Operation,
ctx: KayakContext,
payload: Payload,
next: Next,
) -> BoxFuture<'a, Result<Outcome, KayakError>> {
Box::pin(async move {
if ctx.get::<Tenant>().is_none() {
return Err(KayakError::Unauthorized("no tenant".into()));
}
next.run(operation, ctx, payload).await
})
}
}The router doesn't own a socket. Your HTTP layer authenticates
the request, puts the tenant and caller into a KayakContext,
and hands the router the method, path, query string and body.
This is what it answered in a real run with the middleware
above in place:
GET /v1/tickets?status=open&sort=created_at:desc&limit=500
resolver saw limit=50, filters {status: open}, sort created_at desc
-> 200
GET /v1/tickets?priority=1 -> 400 {"error":"filtering on priority is not allowed"}
GET /v1/tickets?sort=priority -> 400 {"error":"sorting on priority is not allowed"}
GET /v1/tickets (no tenant) -> 401 {"error":"no tenant"}The limit of 500 reached my resolver as 50, the contract's
page ceiling. The undeclared filter and sort never reached it
at all. And when I left out the resolver for close, the
dispatcher wouldn't build:
resource tickets: action close: no resolver registeredAn operation the contract declares with nothing behind it fails the service at startup, instead of turning into a 500 for the first caller who tries it.
The GraphQL face is built from the same dispatcher, so the schema it serves matches the generated SDL by construction:
let schema = kayak::runtime::graphql::schema_builder(&tables, dispatcher)?
.limit_depth(10)
.limit_complexity(500)
.finish()?;Field guardA rule that decides, for each caller and each row, whether a field may be shown. Fields the caller may not see are removed from the answer. cover the columns that some callers may see and others may not. Copal shows who uploaded a file only to that person or to an admin, and the guard is this closure:
pub fn guards() -> kayak::runtime::Guards {
kayak::runtime::Guards::new().guard("owner_or_admin", |ctx, row| {
let Some(principal) = ctx.get::<kayak::runtime::Principal>() else {
return false;
};
if principal.has("admin") {
return true;
}
row.and_then(|r| r.get("created_by"))
.and_then(|v| v.as_str())
.is_some_and(|owner| owner == principal.subject)
})
}The contract names the guard on the field, the dispatcher removes the key from any row the caller shouldn't see, and a caller can't filter or sort on a column they aren't allowed to read. A guard that's declared but never registered, or registered but never used, stops the dispatcher from building, same as a missing resolver.
What it drives
I run Kayak in four services, and each one uses a different amount of it. That turned out to be the most useful property of the design. You can piece-meal features as they are needed.
Describing an API you can't move
The social service behind my PolyConsoleA game platform I'm building, with apps on Android and an online social service behind them. project is hand-written axum,
and its paths are load-bearing because shipped Android builds
call them and can't be rolled back. Moving it onto Kayak's
router would move every path, so it doesn't. Its contract
describes the service that exists instead, with api_prefix
set to an empty string so the routes match what's really
served, and the Differ (Kayak)Kayak's tool that compares two versions of a contract and reports which changes would break the clients that already exist.'s job is to notice when that stops
being true. The service still gets generated clients and a
breaking-change gate without changing a single route.
Antumbra sits at the same layer for a different reason. Its control server publishes read-only OpenAPI and SDL documents for nine resources, generated from the store's real index definitions. Its tables are schemaless, so the contract layer types the wire columns itself and lays them over the indexes the store actually defines. Writes stay on its MCP tools, where the engine's access rules can do reasoning a bare POSTThe kind of HTTP request used to send new data or trigger an action, as opposed to GET, which only reads. couldn't.
A small operator API with typed clients
Penpal hands out expiring links to PenpotAn open-source design and prototyping tool that runs in the browser, similar to Figma. designs. Its operator API is a Kayak contract served through the runtime:
POST /v1/links mint {file_id, mode, ttl_secs?}
GET /v1/links [?mode=view|edit] list, newest first
GET /v1/links/{id} one link
POST /v1/links/{id}/revoke kill it now (idempotent)
GET /v1/links/{id}/redemptions who came, in orderaxumA popular Rust library for building web servers. does the bearer check and hands everything under /v1 to
the RestRouter. The redemptions list is a sub-resource, a
collection that hangs off one link and pages like any other.
The link's storage column was renamed at one point, and a
rename in the contract kept file_id stable on the wire. The
page a guest actually opens, /t/{token}, is left out of the
contract on purpose, because it's a human-facing exchange with
its own grammar and not an API anyone scripts against.
Every face at once
Copal is the reference deployment, and it uses nearly
everything. It has a live GraphQL schema with subscriptions, a
generated REST twin at /v1c, an MCP endpoint where Kayak
writes the tool manifest and Copal's JSON-RPCA simple way for one program to call functions in another by exchanging JSON messages. MCP is built on it. handler calls
the dispatcher directly, and an operator console rendered from
the contract. It adds Rate classA named budget of requests per minute that each caller may spend. When the budget runs out, requests are refused until the next minute. shared across a fleet,
field guards, planner verification in its test suite, and
SurrealDB PERMISSIONS clauses derived from the same
contract, so its scopes and field guards hold a second time
inside the database itself. I've written about Copal
separately in Copal: a file service on one
database.
Tools for agents
The MCP manifest is just another face. Each operation becomes a tool with a typed input schema, and scopes and rate classes ride along as annotations. This is the close action from the helpdesk contract:
{
"name": "ticket_close",
"description": "close on tickets; returns nothing on success.",
"inputSchema": {
"type": "object",
"properties": {
"id": { "type": "string" },
"reason": { "type": "string" }
},
"required": ["id"],
"additionalProperties": false
},
"annotations": {}
}When the agent's call goes through the dispatcher, it gets the same AllowlistA list of what is permitted. Anything not on it is refused., page ceilings, scopes and budgets as every other caller. An agent can't reach anything a key couldn't.
Put together, these are the workloads it's good at:
| Workload | What Kayak does there | Where I use it |
|---|---|---|
| An API that already ships and can't move | Describes it, generates clients, gates changes in CI | polyconsole-social |
| Typed API documents over an existing store | Validates against real indexes, generates OpenAPI and SDL | Antumbra |
| A small internal or operator REST API | Serves it through resolvers you write, behind your HTTP server | Penpal |
| A service with many faces | GraphQL, REST, console and MCP through one dispatcher, with guards and rate classes | Copal |
| Tools for agents | An MCP manifest with scope and rate annotations, enforced like any other caller | Copal |
| Search endpoints | Declared search backings checked against FULLTEXT and vector indexes | Copal |
Starting from a database you already have
Nobody wants to hand-copy every column into a contract and
check every index by eye, so scaffold reads a schema and
writes a contract that already validates:
kayak scaffold --schema schema.json --name helpdesk --out contract.jsonwrote contract.json
1 resources, 3 filters and 2 sorts derived from indexes; narrow to what the API should offer before generatingIt claims less than it could, on purpose. It only claims a
sort where the pinned columns cover everything ahead of it in
the index, because the differ treats a removed sort as
breaking. A claim the scaffold invents would cost a major
version to take back. One it leaves out costs a line to add.
On the helpdesk table it left created_at unclaimed for
exactly that reason. It also exposed internal_note, because
that name doesn't look like a secret. Columns whose names do,
anything with secret, token, password, hash and a few
others, are held back and named on stderrThe channel a command-line program writes warnings and errors to, shown in the terminal separately from its normal output.. Everything else is
your call, which is why the summary line ends by telling you
to narrow it.
Run against Copal's 25 tables, the scaffold exposes 22, derives 39 filters and 28 sorts, and all seven artifacts generate from the result without an edit. The three tables it declines are the ones Copal only ever reaches through a parent: a file's versions, an endpoint's WebhookA message a service sends to an address you choose whenever something happens, so you don't have to keep asking. deliveries, and resumable upload sessions.
What it won't do
Kayak never talks to your database at runtime. Resolvers are your code and read whatever they read, which also makes them the place to make sure a query only selects what the contract exposes.
Actions describe the wire shape, never the behavior. What closing a ticket actually does lives in the service.
Bytes stay with the host. An upload or download face renders as an OpenAPI operation, but there's no GraphQL or MCP shape for streaming a file, and the runtime doesn't stream one.
Authentication is deliberately narrow: no auth, a bearer token, or a named header. OAuthThe standard way one app gets permission to act for you in another, the mechanism behind buttons like Sign in with Google. flows, signed requests and mTLS (mutual TLS)Encrypted connections in which both sides prove who they are with certificates, not only the server. are all real, and none of them is a header a generator can fill in from a constructor argument, so they belong to the service.
It's for SurrealDB and surql-rsThe Rust client library for SurrealDB that Kayak, Copal and Antumbra use. It lets a program declare its database schema as code instead of hand-written queries.. The index rules are the SurrealDB planner's rules, and the schema it checks against is surql-rs table definitions.
And it can't make anyone declare. A query that never mentions
its search index can't be caught by a rule about declared
search indexes. That's the standing limit of any declaration
language, and it's why verify exists beside the static gate.
Current state
Kayak is in pre-release with a scheduled v0.1.0 release. It will be available either through my github, or through crates.io, so for now it's a GitThe version control tool most software is written with. It records every change to a project as a commit, so history can be compared and undone. dependency. It needs Rust 1.90 or newer.
The generated clients are typed where it matters most, in the resources and method names, but their list methods take only a limit and a cursor for now, with no filter or sort parameters, and action inputs are untyped maps.
The built-in rate store is single-process, with fixed one-minute windows that can admit up to double a budget across a window boundary. A fleet needs a shared one, which is a small Trait (Rust)Rust's way of naming a set of functions a type promises to provide, so code can work with any type that provides them. to implement. Copal's runs over SurrealDB. Subscriptions are authorized once, when they open, so revoking access partway through a stream has to be checked inside the resolver.
What I trust is the test suite. There are golden files for every generator, refusal tests that assert on the exact column named, a runtime suite of about 2,400 lines, and a CLI test that compiles the generated Python and parses the generated Go and Rust with the real toolchains wherever they're installed. Four services depend on it today.
Getting it
[dependencies]
oneiriq-kayak = { git = "https://github.com/Oneiriq/kayak", features = ["runtime"] }The crate is oneiriq-kayak and you import it as kayak.
Features are additive:
| Feature | What it adds |
|---|---|
| none | The contract, validation, every generator, the differ and the scaffold |
runtime |
Resolvers, middleware, the dispatcher and the REST router |
graphql |
The live GraphQL schema, on async-graphql |
console |
The operator console, rendered from the contract |
verify |
EXPLAIN checks against a live database |
The CLI installs the usual way:
cargo install --git https://github.com/Oneiriq/kayak --features verifyKayak is licensed under Apache License 2.0A permissive open-source license. You can use, change and redistribute the code, including inside closed-source products and hosted services, as long as you keep the license and copyright notices..
The line it holds
None of the surfaces Kayak writes are hard to write by hand. That was part of the trap. Each one is easy, so each one gets written by whoever is nearby when a feature needs it, and afterwards nothing holds them to each other or to the database. Kayak doesn't write any single surface better than a careful person would. It makes them all come from one place, and it makes the database sign off on the promise before anyone relies on it. The whole idea is to hold one line through a lot of current.
