Concepts
Namespaces, tables, files, Iceberg schemas and snapshots — the data model behind the Data API.
Last updated on
The Data API sits on top of an Apache Iceberg data
lake queried through Trino. Every request is read-only — you browse namespaces and
read schemas with GET, and query or export rows with POST (the body only
carries the query definition). Five building blocks are worth knowing:
Namespace
A namespace is a container identified by a dot-separated path. In this portal
the namespace is usually the publishing organisation's company code — our example
dataset lives in namespace 111958286 (Higienos institutas).
A namespace lists its child namespaces, tables and files, so the whole catalogue is a tree you walk from the root:
GET/namespaces/111958286
{
"success": true,
"data": {
"current_namespace": "111958286",
"child_namespaces": [],
"tables": ["mirusiuju_toksikologinis_rezultatas", "gimimas", "…"],
"files": []
}
}Call GET /namespaces with no path to get
the root list.
Table
A table is an Iceberg table you can do three things with:
Schema
Read the columns, types and nullability.
Query
Read rows as JSON with filter / sort / paging.
Export
Stream the whole table as NDJSON / JSON / CSV / Parquet.
Schema
Before querying, read the table's schema so you know the exact column names and types. The schema is the Iceberg logical schema: every field has a stable numeric id (Iceberg tracks columns by id across renames), a type, and a required (non-nullable) flag.
GET/namespaces/111958286/tables/mirusiuju_toksikologinis_rezultatas/schema
Here is the real schema of our example table:
| ID | Field | Type | Required | Description |
|---|---|---|---|---|
| 1 | _id | string | yes | Unique record identifier. |
| 2 | vda_id | string | — | State Data Agency record id. |
| 3 | mirties_id | string | — | Death-case identifier. |
| 4 | t_rez | string | — | Toxicology test result code. |
| 5 | m_konc | double | — | Measured substance concentration. |
| 6 | mk_vienetai | string | — | Concentration unit code. |
| 7 | pavadinimas | string | — | Substance name. |
| 8 | gr_pavadinimas | string | — | Substance group name. |
| 9 | tipo_pavadinimas | string | — | Substance type name. |
| 10 | lytis | string | — | Sex of the deceased. |
| 11 | mirties_metai | long | — | Year of death. |
| 12 | gyv_tipas | string | — | Residence type (city / rural). |
| 13 | amziaus_grupe | string | — | Age group. |
| 14 | mirties_priezastis | string | — | Cause-of-death group (ICD-10-AM). |
Common Iceberg types you will see: string, long, int, double, boolean,
date, timestamp, and nested struct.
File
A file is a stored object served straight from object storage — a PDF, image
or archive attached to a namespace. The file name is a single basename (no /),
and the response is the raw bytes with the upstream Content-Type.
/namespaces/{namespace}/files/{fileName}
Snapshot
Iceberg keeps a history of table snapshots. Every write creates a new snapshot
with a numeric snapshot_id. The API reads the latest snapshot — queries and
exports always reflect the current table state.
Response envelope
Two response shapes exist across the API — keep the distinction in mind:
Namespace and schema endpoints wrap their payload:
{ "success": true, "data": { "…": "…" } }The query endpoint returns rows plus offset pagination metadata — no success field:
{ "data": [ /* rows */ ], "total": 22143, "total_pages": 4429, "current_page": 1 }Export and file endpoints return raw bytes (NDJSON / JSON / CSV / Excel / Parquet, or the file's content type), not JSON.
Querying tables
Now put the schema to work: select, filter, sort, paginate.
Downloading data
Bulk export as NDJSON / JSON / CSV / Parquet.
See also
How is this guide?