---
title: "Build a spatio-temporal catalog"
description: "Create a custom spatio-temporal dataset catalog with the Python SDK, ingest geospatial metadata, and query it by time, location, and custom fields."
image: "https://tilebox.com/images/tilebox-docs-social-preview.png"
---

> Documentation Index
> Fetch the complete documentation index at: https://tilebox.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Build a spatio-temporal catalog

import { CardGrid } from "@docs/components/ui/card-grid";
import { LinkCard } from "@docs/components/ui/link-card";

Use a spatio-temporal dataset when each datapoint has both a time and a geometry. This is useful for internal imagery catalogs, derived products, ground truth data, regions of interest, and processing outputs that need geospatial lookup.

This guide creates an imagery catalog with the Python SDK. A dataset defines the schema, a collection groups datapoints with that schema, and each datapoint describes one imagery product. The catalog stores metadata and file references; the image files remain in your storage.

## Prerequisites

Install [uv](https://docs.astral.sh/uv/getting-started/installation/) and set your [Tilebox API key](/docs/authentication) in the `TILEBOX_API_KEY` environment variable. If you do not have a Python project yet, run `uv init` first. Add the packages used in this guide, then start JupyterLab:

```bash
uv add tilebox-datasets shapely
uv run --with jupyterlab jupyter lab
```

The `--with jupyterlab` option makes JupyterLab available for this command without adding it to your project dependencies.

Create a notebook with the Python 3 kernel and run the Python snippets below in order. Keep the terminal running while you use JupyterLab.

## Define the catalog schema

The spatio-temporal dataset kind adds `time`, `geometry`, `id`, and `ingestion_time` to the schema. Supply `time` and `geometry` for each datapoint; Tilebox generates `id` and `ingestion_time`.

The `fields` list below defines four custom fields. Python's `str` defines a string field, `float` defines a 64-bit floating-point field, and `Assets` is Tilebox's structured type for file references. Custom fields are optional on individual datapoints; this example supplies all four.

```python title="Python"
from tilebox.datasets import Client
from tilebox.datasets.data.datasets import DatasetKind
from tilebox.datasets.schema import Assets

client = Client()

fields = [
    {
        "name": "product_id",
        "type": str,
        "description": "Stable product or scene identifier from the source catalog.",
        "example_value": "LC08_L2SP_033033_20240808_20240814_02_T1",
        "roles": ["primary_title"],
    },
    {
        "name": "assets",
        "type": Assets,
        "description": "Files associated with the imagery product.",
    },
    {
        "name": "cloud_cover",
        "type": float,
        "description": "Cloud cover percentage for the product footprint.",
        "example_value": "3.2",
        "queryable": True,
    },
    {
        "name": "processing_level",
        "type": str,
        "description": "Processing level of the imagery product.",
        "example_value": "L2_SR",
        "queryable": True,
    },
]
```

`roles: ["primary_title"]` makes `product_id` the datapoint's display title in the Console. It does not enforce uniqueness or replace Tilebox's generated `id`.

`queryable=True` enables server-side filtering on cloud cover and processing level. Field descriptions and example values appear in the generated schema documentation. `example_value` is documentation text, so `"3.2"` is a string here; the ingested cloud cover value will be a float.

The complete schema maps to Python input values as follows:

| Field | Schema type | Python value at ingestion |
| --- | --- | --- |
| `time` | `Timestamp` | Timezone-aware UTC `datetime`. |
| `geometry` | `Geometry` | Shapely geometry in WGS 84 longitude/latitude coordinates (EPSG:4326). |
| `id` | `UUID` | Generated by Tilebox; omit. |
| `ingestion_time` | `Timestamp` | Generated by Tilebox; omit. |
| `product_id` | `string` | `str` containing the source product ID. |
| `assets` | `Assets` | Structured value added by `assets.to_fields()`. |
| `cloud_cover` | `float64` | Cloud cover percentage as a `float`. |
| `processing_level` | `string` | Processing level as a `str`. |

Choose field types and queryable fields before ingestion. Changing or removing existing fields, changing queryability, or adding a queryable field requires all collections in the dataset to be empty. You can add non-queryable fields to a non-empty dataset.

## Create the dataset and collection

Call `create_or_update_dataset` with the dataset kind, code name, field list, and display name. The code name becomes the stable identifier used in SDK calls.

```python title="Python"
dataset = client.create_or_update_dataset(
    kind=DatasetKind.SPATIOTEMPORAL,
    code_name="internal_imagery_catalog",
    fields=fields,
    name="Internal imagery catalog",
)

collection = dataset.get_or_create_collection("landsat_level_2")
```

Repeating these calls with the same definition reuses the dataset and collection. Schema changes remain subject to the restrictions above. Use separate collections to group datapoints by provider, product family, or processing pipeline.

## Prepare datapoints

Prepare one dictionary per imagery product, with keys matching the dataset field names. You can build these dictionaries from an API response, a database query, or a file. No particular source file format is required.

This product references an image and a JSON metadata file. Replace the placeholder metadata and URIs with your own values:

```python title="Python"
from datetime import datetime, timezone
from shapely import box
from tilebox.datasets.assets import (
    Asset, AssetCollection, AssetLocation, MediaType,
)

assets = AssetCollection.from_assets([
    Asset(
        key="image",
        primary=AssetLocation(
            "s3://your-bucket/imagery/example-scene-001.tif"
        ),
        media_type=MediaType.CLOUD_OPTIMIZED_GEOTIFF,
    ),
    Asset(
        key="metadata",
        primary=AssetLocation(
            "s3://your-bucket/imagery/example-scene-001.json"
        ),
        media_type=MediaType.JSON,
    ),
])

records = [
    {
        "time": datetime(2026, 1, 15, 10, 30, tzinfo=timezone.utc),
        "geometry": box(11.2, 46.2, 11.8, 46.8),
        "product_id": "example-scene-001",
        "cloud_cover": 3.2,
        "processing_level": "L2_SR",
        **assets.to_fields(),
    },
]
```

Each `Asset` describes one file, identified by its key (`"image"` or `"metadata"`) within the datapoint. `AssetCollection` groups a datapoint's assets and is separate from the dataset collection. Here, `assets.to_fields()` returns a mapping containing the `assets` field, and `**` merges it into the dictionary.

Set the media type to match your file. Ingestion stores the reference; it does not upload the file or check whether it exists.

## Ingest the catalog

Ingest the prepared records into a collection.

```python title="Python"
collection.ingest(records)
```

## Query by time, location, and custom fields

Find products acquired from January 1 up to, but not including, February 1, 2026, whose footprints intersect `area`, with cloud cover below 10% and processing level `L2_SR`. Tilebox applies these filters on the server.

```python title="Python"
from tilebox.datasets import field

area = box(11.0, 46.0, 12.0, 47.0)

matches = collection.query(
    temporal_extent=("2026-01-01", "2026-02-01"),
    spatial_extent=area,
    filter=(field("cloud_cover") < 10)
    & (field("processing_level") == "L2_SR"),
)

print(matches["product_id"].values.tolist())
```

`query()` returns an `xarray.Dataset` containing matching metadata and asset references, not image pixels. In a new collection, the output is:

```text
['example-scene-001']
```

The [storage client](/docs/datasets/assets-and-storage/read-and-download) accepts these assets directly. Resolve a datapoint's assets with `AssetCollection.from_datapoint(matches.isel(time=0))`. You can download either asset with `storage.download()`, read the JSON file with `storage.read_bytes()`, or open the image with `storage.open_geotiff()`. The referenced files must exist, and you need credentials for private storage; your Tilebox API key does not grant access to the bucket.

## Next steps

Use the [Tilebox Console](/docs/console) or the [CLI](/docs/cli#use-files-and-standard-input) to add Markdown documentation about the dataset's source and use. With the CLI, you can load the schema and documentation from files and keep them in version control.

    Learn the required fields and query behavior.

    Combine queryable field expressions with temporal and spatial filters.

    Load CSV, Parquet, GeoParquet, and NetCDF data before ingestion.

    Download referenced files or read regions from Cloud Optimized GeoTIFFs.

Source: https://tilebox.com/docs/guides/datasets/build-spatiotemporal-catalog/index.mdx
