Copal, a self-hosted file service on SurrealDB

Copal: a file service on one database

Shon Thomas
18 min read

An introduction to Copal, the open-source file service I run on SurrealDB: what happens to a file after you hand it over, why all of the bookkeeping lives in one database, and the work it's good for.

copalsurrealdbstoragerustopen-source

Copal: a file service on one database

Almost everything I build ends up needing somewhere to put files. Antumbra keeps the original bytes of every document it learns from. Penpal is meant to archive the design exports it brokers. PolyConsoleA game platform I'm building, with apps on Android and an online social service behind them.'s players are going to upload things. Each time, the list of wants is the same. Keep the bytes, and don't keep them twice. Check what they actually are. Pull the text out so I can search it. Keep the old versions. Hand out a link that stops working. Tell me when it's done.

The usual way to get that list is to assemble it: an Object storeStorage that keeps files as whole objects under names, like Amazon S3, rather than as folders on one disk. for the bytes, PostgresPostgreSQL, a widely used open-source relational database. for the MetadataData about data. For a file, its name, type, size, owner and history, as opposed to its contents., a Job queueA waiting line of jobs that background workers pick up and process one by one. and WorkerA background process that picks up jobs and does them, apart from the requests people are waiting on. for the processing, a search engine for the text, and something append-only for the Audit trailA permanent record of who did what and when, kept so it can be checked later and cannot be quietly rewritten.. That's five systems to run, and five places for the truth about one file to disagree. The upload finished but the queue message got lost. The search index still has a file the database deleted. A worker crashed halfway through a job and nobody knows which half.

Then, this year, MinIOA popular self-hosted storage server that speaks the Amazon S3 protocol.'s community edition stopped receiving development, security patches and binaries, and most of the replacements hand back a BucketThe top-level container for files in S3-style storage, like a drive that holds folders and files. API (application programming interface)The set of requests one program accepts from another. A web API is how apps, scripts and AI agents ask a service to read or change its data. and nothing else.

Copal is what I built instead, and it's now open source under Apache License 2.0A permissive open-source license. You can use, change and redistribute the code, including inside closed-source products and hosted services, as long as you keep the license and copyright notices..

Most of a file service is bookkeeping, and bookkeeping belongs in one TransactionA group of database changes that succeed or fail together, so the data is never left half updated. database.

Copal keeps file metadata, the search index, the processing journal, access grants, API keys and the audit trail in a single SurrealDBAn open-source database that stores ordinary tables, linked records and nested documents in one engine. Kayak, Copal and Antumbra are all built on it. database. There's no external queue and no workflow engine. Every extra component I could add would be one more thing to operate, and none of them would give me a correctness guarantee the database doesn't already provide.

The name fits the job. Copal is tree resin, the young form of amber, and it holds on to whatever lands in it.

What Copal is

Copal is a Self-hostedSoftware you run on your own machines or your own cloud account, instead of renting it as someone else's service. file service written in RustA programming language known for being fast and for catching whole classes of bugs before a program ever runs., and it ships as one binary. You put files in over RESTThe most common style of web API. Each kind of thing gets a web address, and you read or change it with standard HTTP requests., GraphQLA query language for APIs in which the caller asks for exactly the fields it wants and gets them back in one response., S3Amazon's cloud file storage, and the protocol for storing and fetching files that many other storage systems now speak., tusAn open protocol for uploads that can pause and resume, so a dropped connection doesn't mean starting a large upload over. or MCP (Model Context Protocol)The standard way AI coding agents such as Claude Code connect to outside tools and data.. Copal checks each upload against its DigestA fixed-length fingerprint computed from a file's bytes. Any change to the file changes the digest, so it proves the bytes are intact., stores the bytes once no matter how many times they arrive, keeps every version, runs each upload through a processing pipeline, pulls out the text, and makes it searchable. Then it serves the bytes back under per-file access rules.

The bytes live in a Content addressingNaming data by a cryptographic hash of its bytes instead of by where it is stored. Anyone can serve it, and anyone can check that it matches the name. Blob storeStorage for raw file contents, called blobs, kept separately from the information about them. behind one interface: the local filesystem, S3, GCSGoogle Cloud Storage, Google's equivalent of Amazon S3. or AzureMicrosoft's cloud platform. Its file storage service plays the same role as Amazon S3., optionally Encrypted at restStored on disk in encrypted form, so the raw storage is unreadable to anyone without the key.. Everything else lives in SurrealDB.

One binary, many doors, two planes. The database holds everything about a file except its bytes.

Running it

The default build carries the database engine inside it, so a local instance needs no database server, no ContainerA packaged program bundled with everything it needs to run, so it behaves the same on any machine. and no config file:

Shell
cargo build -p copal-server -p copal-cli

COPAL_BIND=127.0.0.1:8099 COPAL_AUTH_MODE=keys COPAL_ADMIN_TOKEN=local-admin-token COPAL_DB_URL="surrealkv://./data/local/db" COPAL_BLOB_ROOT=./data/local/blobs COPAL_BLOB_ENCRYPTION_KEY=$(printf 'a%.0s' {1..64}) ./target/debug/copal-server

surrealkv:// is SurrealDB running inside the process and persisting to a directory. Point COPAL_DB_URL at ws:// instead and the same binary talks to a SurrealDB server. On first boot it reconciles the SchemaThe declared shape of a database or an API: which tables or types exist, what fields they have, and how they can be looked up. against the empty database and starts listening.

Mint a key for a TenantOne customer, team or workspace whose data is kept separate from everyone else's inside a shared service., create a record, and send it some bytes:

Shell
export COPAL_URL=http://127.0.0.1:8099 COPAL_ADMIN_TOKEN=local-admin-token

TOKEN=$(curl -s -X POST $COPAL_URL/v1/admin/tenants/acme/keys -H "x-copal-admin-token: $COPAL_ADMIN_TOKEN" -H 'content-type: application/json' -d '{"name":"local","scopes":["read","write"]}' | jq -r .token)

ID=$(curl -s -X POST $COPAL_URL/v1/files -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"path":"runbooks/rotation.txt","content_type":"text/plain"}' | jq -r .id)

curl -s -X PUT $COPAL_URL/v1/files/$ID/content -H "authorization: Bearer $TOKEN" --data-binary @rotation.txt

When I ran that for this article, the upload answered like this (trimmed a little):

JSON
{
  "id": "01m3abz35r212xt2tfkva87202",
  "path": "runbooks/rotation.txt",
  "state": "scanning",
  "size": 252,
  "digest": "a0ac806c5553ad3ea9bd1fe18800d269f262501eb5af77dfe35f93ecc9ac5ce8",
  "version_count": 1
}

A moment later the same record read back as ready, with the pipeline's findings written into its metadata:

JSON
"state": "ready",
"metadata": {
  "processing": {
    "sniffed_type": "text/plain",
    "type_matches": true,
    "extracted": true,
    "passages": 1,
    "verdict": "clean"
  }
}

The key that came back looks like ck1.<id>.<secret>, and Copal only stores a hash of the secret. On disk, the bytes sit at objects/a0/ac/<digest> and the file starts with CPE1, the header of Copal's encrypted format, because I gave it a key. The plaintext never touches the filesystem.

The text is searchable straight away:

Shell
curl -s "$COPAL_URL/v1/search?q=rotation" -H "authorization: Bearer $TOKEN"
JSON
{
  "mode": "lexical",
  "items": [
    {
      "file": "01m3abz35r212xt2tfkva87202",
      "passage": 0,
      "excerpt": "Signing key rotation runbook.\nRotate the blob encryption key by...",
      "matches": [[12, 20], [30, 36]]
    }
  ],
  "next_cursor": null
}

The hits come back as passages with character offsets, so a client can highlight them without re-searching. The mode says lexical because I hadn't configured an EmbeddingA list of numbers that captures what a piece of text means, so text about similar things can be found by comparing the numbers. An embedder is the model that produces them. service. Hybrid search is the default, and when Copal can't run the semantic half it answers lexically and says so, instead of quietly returning something different from what you asked for.

The same file is already on every other face. The contract-generated REST twin at /v1c, GraphQL, the MCP EndpointOne specific web address a service answers requests on, such as the one that lists files. and the terminal client all see it:

Shell
curl -s -X POST $COPAL_URL/graphql -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"query":"{ files(limit: 3) { items { id path state } } }"}'
# {"data":{"files":{"items":[{"id":"01m3abz35r212xt2tfkva87202","path":"runbooks/rotation.txt","state":"ready"}]}}}

COPAL_TOKEN=$TOKEN ./target/debug/copalctl search "key rotation" --limit 3

There's also an operator console at /admin/console. It's plain server-rendered HTMLThe language web pages are written in. A page can include scripts that run in the browser of whoever opens it. with no external assets, so it works on an Air-gappedPhysically cut off from other networks, including the internet. network. You log in with any username and the admin token as the password.

What happens to a file

An upload is two requests. The first creates a record, which starts as a draft. The second streams the bytes. Copal hashes them on the way in, writes them to a staging key, and then renames the staged object onto its content address. The record flips to scanning in the same transaction that writes the new version row and links the blob. Then the request returns, and a worker picks up the rest.

An upload, end to end. The request returns once the bytes are safe, and a worker finishes the rest from the journal.

The worker runs six steps in order. It sniffs the real type from the first bytes and compares it with the declared one. It checks the extension policy. If a ClamAVA free, open-source antivirus engine that can scan files as they arrive. daemon is configured it streams the content through it. It extracts the text and splits it into overlapping passages. If an embedding service is configured it embeds them. Finally it flips the record to ready or quarantined and writes the verdict into metadata.processing, which is what you saw above.

Some of that is deliberately done elsewhere. Copal decodes text and JSONA plain-text format for structured data, built from named fields and lists, that almost every programming language can read and write. itself and carries no document ParserA program that reads text or source code and works out its structure. Parsers tend to break when the format they expect changes.. PDF, Office formats and OCR (optical character recognition)Turning pictures of text, such as scanned pages, into real text a computer can search. go to an extractor service with an API shaped like Apache TikaAn open-source tool that pulls the text out of PDFs, Office documents and hundreds of other file formats.'s, because those parsers are large, fast-moving, and historically a rich source of memory-safety bugs. Embeddings come from any /v1/embeddings endpoint shaped like OpenAIThe company behind ChatGPT. Its API format is widely copied, so many tools accept any service shaped like it.'s. Copal runs no models, because InferenceRunning a trained AI model to get an answer, as opposed to training it. means WeightsThe learned numbers inside an AI model. Training adjusts them, and answering a question reads them., a runtime and hardware assumptions that have no business inside a storage service.

The lifecycle of one record. Whether bytes can be served depends on the digest, not on which of these states the record is in.

One rule in that diagram matters more than it looks. Copal serves a file's bytes when the record has a digest and isn't quarantined. The lifecycle state alone never decides. A re-upload in flight doesn't blank out the version that's already there, and a re-upload that fails doesn't take the last good version offline. Tying serving to state would turn every transition into a small outage.

When I uploaded the same bytes a second time, the record went straight to ready as version 2. The digest matched content whose verdict the pipeline had already recorded, so there was nothing to redo, and both versions point at the same blob on disk.

The journal is the queue

The pipeline doesn't need a queue because it runs from a journal in the same database. A run is a row, each step attempt is a row, and a unique index on (run_key, step_key, attempt) means that execution is at-least-once but every step is recorded exactly once. If a worker dies, its LeaseA time-limited claim on a piece of work. If the worker holding it dies, the claim runs out and another worker can take the work. expires, the run goes back to pending, and the next worker replays the journal, skipping every step that already finished. A crash costs one lease timeout.

The rest of the state handling follows the same habit. Every state change carries its preconditions in the WHERE clause of the UPDATE that performs it, and an empty result means the caller lost the race. That makes the database the one place where concurrent writes get ordered, however many Copal instances are running. Every state that can get stuck has an automatic way out:

Stuck on The way out
An upload claimed by an instance that died The lease expires and a sweep marks the file failed, which is retryable
A run claimed by a worker that died The lease expires, the run returns to pending, replay skips finished steps
A file in scanning with no run behind it The stale-scan sweep fails it after a timeout
Staged bytes nobody finished The staging sweep deletes them once they're old enough
Blobs nothing points at Garbage collection recounts from live links, waits out a grace period, then collects

Garbage collectionAutomatically finding and deleting data that nothing refers to anymore. never trusts a counter. It recomputes a blob's references from the live records every time, because counters that get incremented and decremented drift under crashes and races, and a count derived fresh is correct each time it's read.

Custom processing uses the same machinery. A workflow is a list of named activities, each an async function from JSON to JSON, and each step's output becomes the next step's input:

Rust
let registry = FlowRegistry::new()
    .activity("resize", |input| async move {
        // input is the previous step's output (or the run input)
        Ok(serde_json::json!({ "resized": true, "source": input }))
    })
    .workflow("thumbnail", &["resize"], 3);

let state = AppState::new(store, blobs).with_flow(registry);

Activities have to be IdempotentSafe to run more than once: doing it twice has the same effect as doing it once., since a replay after a crash can run one again. For heavier work there's an external transformer seam. The repo ships an ffmpegThe standard open-source tool for converting, cutting and inspecting video and audio files. example, a small PythonA popular, easy-to-read programming language used for everything from small scripts to data science. server with thumbnail, audio, preview and probe recipes. Its probe output walks the normal pipeline and gets indexed like anything else, which means a search for h264 finds videos by CodecThe format audio or video is compressed in, such as H.264 for video..

Search runs over passages rather than whole documents. Lexical searchSearch by the exact words in the text, as opposed to search by meaning. search is BM25A classic formula for ranking text by the words it shares with a query, giving rare words more weight than common ones., and Copal does the scoring itself. Semantic search uses an HNSWHierarchical navigable small world, a popular kind of vector index that finds close matches quickly by hopping through layered graphs of neighbors. index over the passage embeddings, stored at Half and full precisionHow many bits each number in a model uses. Half precision uses 16 bits instead of 32, halving the memory at a small cost in accuracy.. Hybrid, the default, fuses the two rankings with Reciprocal rank fusionA way to merge two ranked lists into one by rewarding items that rank near the top of either list., and an optional RerankerA second model that reads each search result next to the question and reorders the results by how well they actually answer it. can reorder the top results. Filters on where a file sits in its path and on its content type apply inside the database on both halves, so a filtered search never ranks passages it would only throw away.

Text
GET /v1/search?q=signing+key&mode=hybrid&prefix=runbooks/&limit=20
GET /v1/search?q=inspection&facets=content_type,access
GET /v1/files/{id}/text

Embeddings look after themselves. When the configured model changes, a background pass re-embeds every passage that still carries the old model's vectors, and the Vector indexAn index over embeddings that finds the stored items closest in meaning to a query, instead of matching exact words. rebuilds itself when the dimension changes.

Sharing without handing out keys

Most file services hand out Signed URLA link that carries its own permission, so whoever has it can fetch one file without logging in, until it expires., which are awkward to revoke because nothing on the server remembers issuing them. Copal's grants are rows instead. A grant token is cg1.<id>.<secret>, and the database keeps only a hash of the secret, so there's no signing key to rotate or leak. Revoking a grant is one write, and use counts are enforced atomically as part of serving.

A signed URL. The grant is a row, so it can count its uses and be revoked with one write.

In my run I issued a grant with "max_uses": 1. The first request got the file. The second got this:

JSON
{"error":{"kind":"not_found","message":"not found: unknown or unusable grant"}}

Every refusal looks the same, whether the token is malformed, unknown, revoked, expired or used up, so a caller can't probe for which grants exist.

The same idea runs the other way for uploads. Your backend creates the record, decides its path and access level, and asks for an upload grant. The browser or phone then PUTs straight to Copal with nothing but that URLA web address, such as https://example.com/report.pdf.. Upload grants are single-use and capped at 24 hours, and the bytes go through the same size limit, quota and pipeline as any other upload.

For a CDN (content delivery network)A company that keeps copies of websites on servers around the world so pages load quickly. Many large sites depend on the same few CDNs. there's a second family, cg2 edge tokens. They're stateless HMACA signature made with a shared secret key, so anyone holding the same key can check that a message wasn't forged or changed. tokens that an edge worker holding the tenant's edge secret can verify without calling back to the origin. The trade is that you can't revoke one at a time, so they're meant for short lifetimes.

Every file also carries an access level. public files serve anonymously and tell caches to keep them for a year. private and tenant files serve only to their tenant. grant files refuse direct download for everyone, the owner included, so the only way to their bytes is through a URL someone deliberately issued.

One contract behind the API

Copal's API is declared once, as a Kayak Contract (Kayak)In Kayak, the single checked-in declaration of everything an API offers, from which the documents, the clients and the live server are generated.. Kayak checks that declaration against the real indexes and generates the OpenAPIA standard format for describing a REST API. Tools read an OpenAPI document to produce documentation, test tools and client code. document, the GraphQL SDL (Schema Definition Language)The plain-text format that describes a GraphQL API: its types, their fields, and the queries and changes it accepts., the MCP manifest and clients in four languages. It also runs the GraphQL, MCP, /v1c and console faces through one Dispatcher (Kayak)The part of Kayak's runtime that checks each request against the contract, applies limits and permissions, and only then calls your code. that enforces scopes, rate budgets and Field guardA rule that decides, for each caller and each row, whether a field may be shown. Fields the caller may not see are removed from the answer.. The hand-written /v1 routes stay canonical for the things a generic router doesn't do yet, such as streaming bytes, ranges and conditional requests. A parity test holds the two REST faces to the same answers.

That matters most for agents. tools/list returned 22 tools in my run, from files_list and file_create through search and run_start. An agent's whole ingest loop is three calls: file_create, file_issue_upload_url, and an HTTPThe protocol web browsers, apps and servers use to request and send data over the web. PUT of the bytes to the grant URL. It goes through the same scopes, budgets and guards as a script with the same key would. What an agent may do is what its key may do. Keys can also belong to principals of kind agent, which get their own budgets and their own line in the audit trail.

An agent searching through MCP. The tool call goes through the same checks as any other caller holding the same key.

Coming from MinIO

The S3 gatewayCopal's door for tools that speak the S3 protocol, so the same commands that work with Amazon S3 or MinIO work with Copal. takes the tools people already have. The bucket is the tenant and the key is the path, so a migration is one mirror:

Shell
mc alias set old   http://minio:9000 <old key>   <old secret>
mc alias set copal http://copal:9000 <access key> <secret>
mc mirror old/data copal/acme
mc diff old/data copal/acme          # empty output means every object arrived

In the recorded run in the migration guide, 80 MiB (mebibyte)1,048,576 bytes, a little more than a million. came across in 13 seconds, a second pass copied nothing, and an 80 MiB Multipart uploadUploading a large file as separately sent pieces that the server joins at the end, so one failed piece can be retried on its own. object round-tripped with a matching SHA-256A widely used cryptographic hash function that produces a 256-bit fingerprint of any data.. Every file that arrives that way gets scanned, DedupeStoring identical content only once, however many times it arrives or under however many names., versioned and indexed like any other upload, so the mirrored bucket can answer search queries, which a plain object store can't.

The Conformance suiteA set of tests that checks whether a system behaves the way a standard, or the existing tools that speak it, expect. runs stock clients against the S3 gateway, and anyone can run it on their own hardware:

Client Checks passed
aws CLI 9 of 9
MinIO mc 8 of 8
rclone 5 of 5
MCP over curl 4 of 4

The S3 gateway is a door onto Copal's storage and not a full S3 implementation. Tagging, bucket policies, lifecycle rules, ACL (access control list)A list attached to a file or resource that says exactly who may read or change it., the S3 versioning API and Object Lock all answer 501 NotImplemented. ETagA short tag a server attaches to a file so a client can tell whether the file has changed since it last looked. are SHA-256 rather than MD5An older hash function, still used as a quick file fingerprint but no longer considered secure., which means aws s3 sync sees every object as changed.

What it drives

These are the jobs Copal is built for, and the ones I use it for or am building toward:

Workload What Copal gives it
A document store behind an AI memory system Idempotent creates, versions per re-ingest, extracted text, provenance by digest
A searchable archive or RAG source Passage-level lexical, semantic and hybrid search, with facets and an optional reranker
Uploads from browsers and phones Single-use upload grants, so the client never holds a key
Media behind a CDN Immutable public caching, edge tokens, renditions, an ffmpeg transformer seam
A MinIO replacement An S3 gateway that stock tools accept, plus everything above for each migrated object
Compliance-sensitive storage Retention, legal hold, per-residency encryption keys, an audit trail the database refuses to rewrite
Multi-tenant SaaS backends Per-tenant keys, quotas and storage residencies, with tenancy optionally enforced a second time by the database
Event-driven integrations Signed webhooks, a replayable change feed, GraphQL subscriptions
Tools for agents 22 MCP tools under the same scopes and budgets as every other caller

The first row is live today. Antumbra archives the original bytes of every document it ingests into Copal, and every passage it stores carries the Copal file id and digest it came from. The whole client is two calls:

Rust
let mut body = json!({
    "path": document_path(&title, hash),
    "content_type": "text/plain",
    "idempotency_key": idempotency_key(hash),
    "metadata": { "title": title, "workspace": workspace },
});
if let Some(src) = source {
    body["metadata"]["source"] = json!(src);
}
let created = transport.post_json(&format!("{base}/v1/files"), &credential, &body)?;
let file_id = created
    .get("id")
    .and_then(Value::as_str)
    .ok_or_else(|| AntumbraError::other("copal create response missing `id`"))?;
let uploaded = transport.put_bytes(
    &format!("{base}/v1/files/{file_id}/content"),
    &credential,
    "text/plain",
    content.as_bytes(),
)?;

The idempotency key is derived from the workspace and title, so ingesting the same document again makes a new version of the same Copal file rather than a second file.

My own instance runs in k3sA lightweight version of Kubernetes, the system that runs and restarts containerized services across a group of machines. on my homelab as a single ReplicaOne running copy of a service. Several replicas can share the load or take over if one fails. with a 2 GiB (gibibyte)1,073,741,824 bytes, a little more than a billion. memory limit, against its own SurrealDB. Its embeddings come from a Jetson Orin NanoA small NVIDIA computer with a built-in GPU, made for running AI models on very little power. running llama.cppAn open-source program for running AI models efficiently on ordinary hardware, from laptops to small boards. on the same network.

It scales down further than that and up a little past it. The same binary runs in three shapes:

Three topologies, one binary. The embedded shape has nothing else to operate, and two instances need no sticky sessions.

The two-instance shape is tested, not just drawn. COPAL_HA=1 ./conformance/run.sh runs the whole conformance suite through nginxA widely used web server, often placed in front of other services to spread requests between them. Round-robinHanding each new request to the next server in turn. with no Sticky sessionsSending every request from one user to the same server. A service that doesn't need them lets any server answer any request.. The instances need a shared blob root, the same keys and one database. Sweeps elect a leader through a database lease, and rate budgets can be shared across the fleet through the database too.

Some numbers

These come from my own machine: an i9-12900KSA high-end Intel desktop processor. with 64 GiB of RAMA computer's working memory, which holds whatever it is actively using. It is much faster than disk and is cleared when the power goes off. and NVMeA fast kind of solid-state storage that connects directly to the computer's processor. storage, with in-memory metadata, filesystem blobs, and encryption turned on. They describe the shape of the system rather than a tuned deployment's ceiling.

Measurement 512 MiB 1 GiB 2 GiB
Streamed single PUT 1.6 s 3.4 s 6.4 s
Ranged 1 MiB read, start / middle / end 2 / 2 / 1 ms 2 / 3 / 2 ms 2 / 2 / 2 ms
Full sequential read 0.9 s 1.8 s 3.7 s
Peak memory, PUT and every read 49 MiB 49 MiB 50 MiB

Sealed uploads run at about 300 MiB/s, against 515 MiB/s for the same 1 GiB upload unencrypted. Ranged readReading just part of a file instead of the whole thing, which is how a video player jumps to the middle of a video. into an encrypted object stay around 2 ms (millisecond)A thousandth of a second. wherever they land, because the format seals in 64 KiB (kibibyte)1,024 bytes, a little more than a thousand. frames and a read only opens the frames it needs. Memory stays flat as files grow, with one exception: the malware scan pass holds about three times the object, which is worth knowing before you send ClamAV a 2 GiB video.

The embedded engine is also quick. A point read of one file record takes about 109 µs (microsecond)A millionth of a second. at the median when SurrealDB runs inside the process, against about 606 µs over a WebSocketA long-lived, two-way connection between a program and a server, so messages can flow either way without a new request each time. to a server on the same host.

What it won't do

It won't parse documents in-process or run models. Both are seams to services you run next to it.

It isn't a full S3 implementation. The S3 gateway covers what migration and everyday sync tools need, and stops there.

It won't render content that can run script. HTML, SVGA web image format for drawings, written as text. Because it can also carry scripts, services treat it as potentially active content. and XMLA text format for structured data built from nested tags, older and wordier than JSON. are always served as attachments.

It doesn't do multi-region writes. Replication is designed but not built, and for now a second region means a second deployment.

It keeps no write-ahead archive of its own, so point-in-time recovery comes from backing up the database and the blob store.

It doesn't have its own operator sign-in. Operator identity comes from a ProxyA server that sits in front of another and passes requests through to it, often handling tasks such as login on its behalf. in front of it, such as oauth2-proxy, Authelia or Cloudflare Access.

Current state of Copal

Copal hasn't cut a 0.1.0 release. CI (continuous integration)Automated checks that build and test every proposed change before it is merged, so problems are caught before they ship. publishes a container image to ghcr.io/oneiriq/copal on every push to main, and the code is what I run.

The defaults are open for development on purpose. The default auth mode trusts an x-copal-tenant header, and it flips to API keys at 1.0. Before exposing an instance, turn on keys mode with an admin token, put the admin surface on its own listener, terminate TLS (Transport Layer Security)The encryption that protects web traffic, shown by the padlock in your browser's address bar. in front of it, and give it a scoped database user. The operations guide walks through all of it.

The generated client SDK (software development kit)A ready-made library that makes a service easy to use from one programming language. exist in Rust, TypeScriptJavaScript with type annotations added, widely used for web apps and servers. The types catch many mistakes before the code runs., Python and Go, but they aren't published to any registry yet. They authenticate with the development tenant header only, and they don't move bytes. Uploads and downloads are plain HTTP, or a PUT to an upload grant.

Search has two limits worth knowing. BM25 ranks a bounded window of matches, so a query that matches more than a few hundred passages ranks that window rather than every match. The search cursor is best-effort, because rankings shift as content changes.

What I lean on is the test suite. The whole service integration-tests against an in-memory SurrealDB engine with no containers involved, across 44 integration test files in the server alone, alongside the conformance suite and the BenchmarkA fixed set of test questions or tasks used to measure a system and compare it with others. above.

Getting it

Build from source with the commands in Running it, or pull the container image. The development loop is three commands:

Shell
docker compose up -d          # SurrealDB v3, for the server-backed shape
cargo test                    # unit and in-memory round-trip tests, no server needed
cargo run -p copal-server

Copal is licensed under Apache-2.0.

Retrospect

The part of Copal I'm happiest with is how little there is to run. A file's bytes are in one place, and everything anyone knows about them (versions, verdicts, passages, grants, events, the audit trail) is in one database that can answer for all of it in a single transaction. When something goes wrong, there's one place to look. When something gets stuck, there's a sweep that knows the way out. It holds on to what you give it, and it can tell you what it's holding.