Skip to content

Build a spatio-temporal catalog

Create a custom spatio-temporal dataset catalog with the Python SDK, ingest geospatial metadata, and query it by time, location, and custom fields.

Use a spatio-temporal dataset when each datapoint has both a time and a geometry. This is useful for internal imagery catalogs, derived products, ground truth data, regions of interest, and processing outputs that need geospatial lookup.

This guide creates an imagery catalog with the Python SDK. A dataset defines the schema, a collection groups datapoints with that schema, and each datapoint describes one imagery product. The catalog stores metadata and file references; the image files remain in your storage.

Prerequisites

Install uv and set your Tilebox API key in the TILEBOX_API_KEY environment variable. If you do not have a Python project yet, run uv init first. Add the packages used in this guide, then start JupyterLab:

bash
uv add tilebox-datasets shapely
uv run --with jupyterlab jupyter lab

The --with jupyterlab option makes JupyterLab available for this command without adding it to your project dependencies.

Create a notebook with the Python 3 kernel and run the Python snippets below in order. Keep the terminal running while you use JupyterLab.

Define the catalog schema

The spatio-temporal dataset kind adds time, geometry, id, and ingestion_time to the schema. Supply time and geometry for each datapoint; Tilebox generates id and ingestion_time.

The fields list below defines four custom fields. Python’s str defines a string field, float defines a 64-bit floating-point field, and Assets is Tilebox’s structured type for file references. Custom fields are optional on individual datapoints; this example supplies all four.

Python
from tilebox.datasets import Client
from tilebox.datasets.data.datasets import DatasetKind
from tilebox.datasets.schema import Assets

client = Client()

fields = [
    {
        "name": "product_id",
        "type": str,
        "description": "Stable product or scene identifier from the source catalog.",
        "example_value": "LC08_L2SP_033033_20240808_20240814_02_T1",
        "roles": ["primary_title"],
    },
    {
        "name": "assets",
        "type": Assets,
        "description": "Files associated with the imagery product.",
    },
    {
        "name": "cloud_cover",
        "type": float,
        "description": "Cloud cover percentage for the product footprint.",
        "example_value": "3.2",
        "queryable": True,
    },
    {
        "name": "processing_level",
        "type": str,
        "description": "Processing level of the imagery product.",
        "example_value": "L2_SR",
        "queryable": True,
    },
]

roles: ["primary_title"] makes product_id the datapoint’s display title in the Console. It does not enforce uniqueness or replace Tilebox’s generated id.

queryable=True enables server-side filtering on cloud cover and processing level. Field descriptions and example values appear in the generated schema documentation. example_value is documentation text, so "3.2" is a string here; the ingested cloud cover value will be a float.

The complete schema maps to Python input values as follows:

Field Schema type Python value at ingestion
time Timestamp Timezone-aware UTC datetime.
geometry Geometry Shapely geometry in WGS 84 longitude/latitude coordinates (EPSG:4326).
id UUID Generated by Tilebox; omit.
ingestion_time Timestamp Generated by Tilebox; omit.
product_id string str containing the source product ID.
assets Assets Structured value added by assets.to_fields().
cloud_cover float64 Cloud cover percentage as a float.
processing_level string Processing level as a str.

Choose field types and queryable fields before ingestion. Changing or removing existing fields, changing queryability, or adding a queryable field requires all collections in the dataset to be empty. You can add non-queryable fields to a non-empty dataset.

Create the dataset and collection

Call create_or_update_dataset with the dataset kind, code name, field list, and display name. The code name becomes the stable identifier used in SDK calls.

Python
dataset = client.create_or_update_dataset(
    kind=DatasetKind.SPATIOTEMPORAL,
    code_name="internal_imagery_catalog",
    fields=fields,
    name="Internal imagery catalog",
)

collection = dataset.get_or_create_collection("landsat_level_2")

Repeating these calls with the same definition reuses the dataset and collection. Schema changes remain subject to the restrictions above. Use separate collections to group datapoints by provider, product family, or processing pipeline.

Prepare datapoints

Prepare one dictionary per imagery product, with keys matching the dataset field names. You can build these dictionaries from an API response, a database query, or a file. No particular source file format is required.

This product references an image and a JSON metadata file. Replace the placeholder metadata and URIs with your own values:

Python
from datetime import datetime, timezone
from shapely import box
from tilebox.datasets.assets import (
    Asset, AssetCollection, AssetLocation, MediaType,
)

assets = AssetCollection.from_assets([
    Asset(
        key="image",
        primary=AssetLocation(
            "s3://your-bucket/imagery/example-scene-001.tif"
        ),
        media_type=MediaType.CLOUD_OPTIMIZED_GEOTIFF,
    ),
    Asset(
        key="metadata",
        primary=AssetLocation(
            "s3://your-bucket/imagery/example-scene-001.json"
        ),
        media_type=MediaType.JSON,
    ),
])

records = [
    {
        "time": datetime(2026, 1, 15, 10, 30, tzinfo=timezone.utc),
        "geometry": box(11.2, 46.2, 11.8, 46.8),
        "product_id": "example-scene-001",
        "cloud_cover": 3.2,
        "processing_level": "L2_SR",
        **assets.to_fields(),
    },
]

Each Asset describes one file, identified by its key ("image" or "metadata") within the datapoint. AssetCollection groups a datapoint’s assets and is separate from the dataset collection. Here, assets.to_fields() returns a mapping containing the assets field, and ** merges it into the dictionary.

Set the media type to match your file. Ingestion stores the reference; it does not upload the file or check whether it exists.

Ingest the catalog

Ingest the prepared records into a collection.

Python
collection.ingest(records)

Query by time, location, and custom fields

Find products acquired from January 1 up to, but not including, February 1, 2026, whose footprints intersect area, with cloud cover below 10% and processing level L2_SR. Tilebox applies these filters on the server.

Python
from tilebox.datasets import field

area = box(11.0, 46.0, 12.0, 47.0)

matches = collection.query(
    temporal_extent=("2026-01-01", "2026-02-01"),
    spatial_extent=area,
    filter=(field("cloud_cover") < 10)
    & (field("processing_level") == "L2_SR"),
)

print(matches["product_id"].values.tolist())

query() returns an xarray.Dataset containing matching metadata and asset references, not image pixels. In a new collection, the output is:

text
['example-scene-001']

The storage client accepts these assets directly. Resolve a datapoint’s assets with AssetCollection.from_datapoint(matches.isel(time=0)). You can download either asset with storage.download(), read the JSON file with storage.read_bytes(), or open the image with storage.open_geotiff(). The referenced files must exist, and you need credentials for private storage; your Tilebox API key does not grant access to the bucket.

Next steps

Use the Tilebox Console or the CLI to add Markdown documentation about the dataset’s source and use. With the CLI, you can load the schema and documentation from files and keep them in version control.

Type to search…