Getting started

Concepts

Namespaces, tables, files, Iceberg schemas and snapshots — the data model behind the Data API.

Last updated on

The Data API sits on top of an Apache Iceberg data lake queried through Trino. Every request is read-only — you browse namespaces and read schemas with GET, and query or export rows with POST (the body only carries the query definition). Five building blocks are worth knowing:

Namespace

A namespace is a container identified by a dot-separated path. In this portal the namespace is usually the publishing organisation's company code — our example dataset lives in namespace 111958286 (Higienos institutas).

A namespace lists its child namespaces, tables and files, so the whole catalogue is a tree you walk from the root:

GET/namespaces/111958286
{
  "success": true,
  "data": {
    "current_namespace": "111958286",
    "child_namespaces": [],
    "tables": ["mirusiuju_toksikologinis_rezultatas", "gimimas", "…"],
    "files": []
  }
}

Call GET /namespaces with no path to get the root list.

Table

A table is an Iceberg table you can do three things with:

Schema

Before querying, read the table's schema so you know the exact column names and types. The schema is the Iceberg logical schema: every field has a stable numeric id (Iceberg tracks columns by id across renames), a type, and a required (non-nullable) flag.

GET/namespaces/111958286/tables/mirusiuju_toksikologinis_rezultatas/schema

Here is the real schema of our example table:

mirusiuju_toksikologinis_rezultatas — 14 fields
IDFieldTypeRequiredDescription
1_idstringyesUnique record identifier.
2vda_idstring—State Data Agency record id.
3mirties_idstring—Death-case identifier.
4t_rezstring—Toxicology test result code.
5m_koncdouble—Measured substance concentration.
6mk_vienetaistring—Concentration unit code.
7pavadinimasstring—Substance name.
8gr_pavadinimasstring—Substance group name.
9tipo_pavadinimasstring—Substance type name.
10lytisstring—Sex of the deceased.
11mirties_metailong—Year of death.
12gyv_tipasstring—Residence type (city / rural).
13amziaus_grupestring—Age group.
14mirties_priezastisstring—Cause-of-death group (ICD-10-AM).

Common Iceberg types you will see: string, long, int, double, boolean, date, timestamp, and nested struct.

File

A file is a stored object served straight from object storage — a PDF, image or archive attached to a namespace. The file name is a single basename (no /), and the response is the raw bytes with the upstream Content-Type.

GET/namespaces/{namespace}/files/{fileName}

Snapshot

Iceberg keeps a history of table snapshots. Every write creates a new snapshot with a numeric snapshot_id. The API reads the latest snapshot — queries and exports always reflect the current table state.

Response envelope

Two response shapes exist across the API — keep the distinction in mind:

Namespace and schema endpoints wrap their payload:

{ "success": true, "data": { "…": "…" } }

The query endpoint returns rows plus offset pagination metadata — no success field:

{ "data": [ /* rows */ ], "total": 22143, "total_pages": 4429, "current_page": 1 }

Export and file endpoints return raw bytes (NDJSON / JSON / CSV / Excel / Parquet, or the file's content type), not JSON.

See also

How is this guide?