Build a spatio-temporal catalog
Create a custom spatio-temporal dataset catalog with the Python SDK, ingest geospatial metadata, and query it by time, location, and custom fields.
Use a spatio-temporal dataset when each datapoint has both a time and a geometry. This is useful for internal imagery catalogs, derived products, ground truth data, regions of interest, and processing outputs that need geospatial lookup.
This guide creates an imagery catalog with the Python SDK. A dataset defines the schema, a collection groups datapoints with that schema, and each datapoint describes one imagery product. The catalog stores metadata and file references; the image files remain in your storage.
Prerequisites
Install uv and set your Tilebox API key in the TILEBOX_API_KEY environment variable. If you do not have a Python project yet, run uv init first. Add the packages used in this guide, then start JupyterLab:
uv add tilebox-datasets shapely
uv run --with jupyterlab jupyter labThe --with jupyterlab option makes JupyterLab available for this command without adding it to your project dependencies.
Create a notebook with the Python 3 kernel and run the Python snippets below in order. Keep the terminal running while you use JupyterLab.
Define the catalog schema
The spatio-temporal dataset kind adds time, geometry, id, and ingestion_time to the schema. Supply time and geometry for each datapoint; Tilebox generates id and ingestion_time.
The fields list below defines four custom fields. Python’s str defines a string field, float defines a 64-bit floating-point field, and Assets is Tilebox’s structured type for file references. Custom fields are optional on individual datapoints; this example supplies all four.
from tilebox.datasets import Client
from tilebox.datasets.data.datasets import DatasetKind
from tilebox.datasets.schema import Assets
client = Client()
fields = [
{
"name": "product_id",
"type": str,
"description": "Stable product or scene identifier from the source catalog.",
"example_value": "LC08_L2SP_033033_20240808_20240814_02_T1",
"roles": ["primary_title"],
},
{
"name": "assets",
"type": Assets,
"description": "Files associated with the imagery product.",
},
{
"name": "cloud_cover",
"type": float,
"description": "Cloud cover percentage for the product footprint.",
"example_value": "3.2",
"queryable": True,
},
{
"name": "processing_level",
"type": str,
"description": "Processing level of the imagery product.",
"example_value": "L2_SR",
"queryable": True,
},
]roles: ["primary_title"] makes product_id the datapoint’s display title in the Console. It does not enforce uniqueness or replace Tilebox’s generated id.
queryable=True enables server-side filtering on cloud cover and processing level. Field descriptions and example values appear in the generated schema documentation. example_value is documentation text, so "3.2" is a string here; the ingested cloud cover value will be a float.
The complete schema maps to Python input values as follows:
| Field | Schema type | Python value at ingestion |
|---|---|---|
time |
Timestamp |
Timezone-aware UTC datetime. |
geometry |
Geometry |
Shapely geometry in WGS 84 longitude/latitude coordinates (EPSG:4326). |
id |
UUID |
Generated by Tilebox; omit. |
ingestion_time |
Timestamp |
Generated by Tilebox; omit. |
product_id |
string |
str containing the source product ID. |
assets |
Assets |
Structured value added by assets.to_fields(). |
cloud_cover |
float64 |
Cloud cover percentage as a float. |
processing_level |
string |
Processing level as a str. |
Choose field types and queryable fields before ingestion. Changing or removing existing fields, changing queryability, or adding a queryable field requires all collections in the dataset to be empty. You can add non-queryable fields to a non-empty dataset.
Create the dataset and collection
Call create_or_update_dataset with the dataset kind, code name, field list, and display name. The code name becomes the stable identifier used in SDK calls.
dataset = client.create_or_update_dataset(
kind=DatasetKind.SPATIOTEMPORAL,
code_name="internal_imagery_catalog",
fields=fields,
name="Internal imagery catalog",
)
collection = dataset.get_or_create_collection("landsat_level_2")Repeating these calls with the same definition reuses the dataset and collection. Schema changes remain subject to the restrictions above. Use separate collections to group datapoints by provider, product family, or processing pipeline.
Prepare datapoints
Prepare one dictionary per imagery product, with keys matching the dataset field names. You can build these dictionaries from an API response, a database query, or a file. No particular source file format is required.
This product references an image and a JSON metadata file. Replace the placeholder metadata and URIs with your own values:
from datetime import datetime, timezone
from shapely import box
from tilebox.datasets.assets import (
Asset, AssetCollection, AssetLocation, MediaType,
)
assets = AssetCollection.from_assets([
Asset(
key="image",
primary=AssetLocation(
"s3://your-bucket/imagery/example-scene-001.tif"
),
media_type=MediaType.CLOUD_OPTIMIZED_GEOTIFF,
),
Asset(
key="metadata",
primary=AssetLocation(
"s3://your-bucket/imagery/example-scene-001.json"
),
media_type=MediaType.JSON,
),
])
records = [
{
"time": datetime(2026, 1, 15, 10, 30, tzinfo=timezone.utc),
"geometry": box(11.2, 46.2, 11.8, 46.8),
"product_id": "example-scene-001",
"cloud_cover": 3.2,
"processing_level": "L2_SR",
**assets.to_fields(),
},
]Each Asset describes one file, identified by its key ("image" or "metadata") within the datapoint. AssetCollection groups a datapoint’s assets and is separate from the dataset collection. Here, assets.to_fields() returns a mapping containing the assets field, and ** merges it into the dictionary.
Set the media type to match your file. Ingestion stores the reference; it does not upload the file or check whether it exists.
Ingest the catalog
Ingest the prepared records into a collection.
collection.ingest(records)Query by time, location, and custom fields
Find products acquired from January 1 up to, but not including, February 1, 2026, whose footprints intersect area, with cloud cover below 10% and processing level L2_SR. Tilebox applies these filters on the server.
from tilebox.datasets import field
area = box(11.0, 46.0, 12.0, 47.0)
matches = collection.query(
temporal_extent=("2026-01-01", "2026-02-01"),
spatial_extent=area,
filter=(field("cloud_cover") < 10)
& (field("processing_level") == "L2_SR"),
)
print(matches["product_id"].values.tolist())query() returns an xarray.Dataset containing matching metadata and asset references, not image pixels. In a new collection, the output is:
['example-scene-001']The storage client accepts these assets directly. Resolve a datapoint’s assets with AssetCollection.from_datapoint(matches.isel(time=0)). You can download either asset with storage.download(), read the JSON file with storage.read_bytes(), or open the image with storage.open_geotiff(). The referenced files must exist, and you need credentials for private storage; your Tilebox API key does not grant access to the bucket.
Next steps
Use the Tilebox Console or the CLI to add Markdown documentation about the dataset’s source and use. With the CLI, you can load the schema and documentation from files and keep them in version control.
Spatio-temporal datasets
Learn the required fields and query behavior.
Filter by custom fields
Combine queryable field expressions with temporal and spatial filters.
Ingest from common file formats
Load CSV, Parquet, GeoParquet, and NetCDF data before ingestion.
Read and download assets
Download referenced files or read regions from Cloud Optimized GeoTIFFs.