# Deployment

From a laptop to a cluster. Read the honesty notes — two of them change how you size
things.

## Single machine

Nothing to deploy. `python main.py` for the desktop shell, `run_pipeline.py` for
headless processing. Data stays on disk, nothing calls out.

## Docker Compose

```bash
cd infrastructure/docker
docker compose up
```

Four services:

| Service | Role |
|---|---|
| `postgis` | PostgreSQL with the PostGIS extension |
| `minio` | S3-compatible object storage |
| `api` | The FastAPI application |
| `worker` | Processing jobs |

This is the intended shape for a small team on one host.

## Kubernetes (Helm)

```bash
helm install odk infrastructure/helm/opendronekit \
  --set secrets.odkSecretKey=<value> \
  --set secrets.postgresPassword=<value>
```

### There are no default credentials, and rendering fails without them

`values.yaml` ships every secret empty, and the chart's helpers call `fail` when one is
missing. This is deliberate: "change this in production" is advice nobody reads in time,
and a default that works is a default that reaches production. A chart that lints
cleanly and ships a working default credential is worse than one that refuses to render.

If you manage secrets externally, set `secrets.existingSecret` instead. Refusing
defaults is only reasonable because there is a real alternative to them.

### What the chart provides

- API deployment, with **readiness on `/health/ready`** and liveness on `/health/live`.
- Worker deployment with a **memory limit**. Photogrammetry is memory-bound before it
  is CPU-bound, and an unlimited worker is OOM-killed mid-reconstruction — losing hours
  rather than degrading.
- PostGIS as a **StatefulSet with a volume claim**. Survey data is the one thing here
  that cannot be regenerated by re-running a job, so it never sits on an `emptyDir`.

### What the chart has not been validated against

`tests/test_helm_chart.py` performs structural checks only. Without the `helm` binary it
cannot render the Go templates, so **nothing proves the chart produces valid Kubernetes
objects**. `helm lint` and `helm template` against a real cluster are what move
`inf.k8s` from implemented to verified. Run them before you rely on this.

## Storage

Local filesystem and S3-compatible backends sit behind one interface. Keys cannot escape
the storage root, and an unknown backend is **refused rather than silently falling back
to local disk** — a fallback that writes survey data somewhere unexpected is worse than
a startup failure.

## Database and spatial storage

Set `ODK_DATABASE_URL` to a PostgreSQL instance for multi-user deployment. `init_db()`
creates the PostGIS extension where the database permits it, then applies the spatial
migration.

**Geometry is stored as GeoJSON text and mirrored into GIST-indexed PostGIS geometry
columns** on `assets`, `defects`, `measurements` and `annotations`. A trigger keeps the
mirror in step, so no application path can forget it — writes arrive through several
routers and a plugin SDK, and one that skipped the mirror would leave a row findable by
text and invisible to spatial queries.

The text column remains the **source of truth**. That is deliberate: SQLite is a
supported backend and cannot hold a geometry type, so making the native column
authoritative would fork the schema and give the two backends different answers to the
same question. The `geom` column is a derived index.

A row whose GeoJSON will not parse is still **stored**, with `geom` NULL. It stays
visible to every non-spatial query and falls back to the text path. Rejecting the write
would lose data in order to gain an index.

Check what you have:

```bash
curl localhost:8000/health | jq '.spatial'
```

```json
{"backend": "postgresql", "postgis": true,
 "geometry_storage": "geojson_text_with_native_mirror",
 "native_geometry_columns": true,
 "indexed_tables": ["annotations", "assets", "defects", "measurements"]}
```

The report reads the columns, not the extension. It used to say `native_geometry`
whenever PostGIS answered a version query while every column was still Text — which
would have had you sizing a query-heavy workload around indexes that did not exist.

On SQLite, `native_geometry_columns` is `false` and spatial filtering happens in Python.
That is the honest answer for a development database, and the API says so rather than
letting you assume otherwise.

### Verifying it yourself

```bash
docker run -d --name odk-postgis -e POSTGRES_PASSWORD=odk -e POSTGRES_DB=odk     -p 55432:5432 postgis/postgis:16-3.4
ODK_TEST_POSTGIS=postgresql+psycopg://postgres:odk@127.0.0.1:55432/odk python -m pytest     tests/test_postgis_spatial.py
```

Those tests skip without a live instance rather than passing against SQLite, because a
spatial test that never touches PostGIS proves nothing about spatial behaviour.

## Observability

Prometheus metrics at `/metrics`. No external telemetry is sent anywhere by default, and
that is a property the project tests for rather than a promise in a document
(`inf.offline_first`).

## What CI checks, and why it is four jobs

`tests` is the fast gate. `spatial` and `sitl` exist because a test that skips when its
dependency is absent reports success while checking nothing — right on a laptop, a lie in
CI — so both assert their tests actually **ran** rather than skipped.

`status` runs last and needs the other two, because feature status is computed from
passing tests and a skip counts as no evidence. That makes `fl.sitl` unearnable anywhere
except a machine with ArduPilot, so the `sitl` job publishes its junit report and `status`
merges it with `tools/feature_status.py --extra-report`. A failure in the merged report
still counts as a failure: an outside run can promote a row by passing, never by being
quieter than the local one. A missing or empty report fails the job.

The practical consequence for anyone reading a status: **`verified` means tests named by
that row passed in a run that happened**, not that someone decided the feature was done.

## SITL

`infrastructure/docker/Dockerfile.sitl` builds ArduPilot Copter-4.5.7 with the test
dependencies, so flight code can be verified against a real autopilot without Linux
hardware. See `docs/SITL.md`. The build is pinned to a release tag on purpose: a
simulator whose behaviour moves under the tests turns a red run into an unanswerable
question.
