Skip to content

Ingesting Data

Populate your dataset collections by defining schemas, preparing structured data points, and submitting them efficiently to Tilebox for storage and querying.

Check out the examples below for common scenarios of ingesting data into a collection.

Dataset schema

Tilebox Datasets are strongly typed. This means you can only ingest data that matches the schema of a dataset. The schema is defined during dataset creation time.

The examples on this page assume that you have access to a Timeseries dataset that has the following schema:

MyCustomDataset schema

Field name Type Description
time Timestamp Timestamp of the data point. Required by the Timeseries dataset type.
id UUID Auto-generated UUID for each datapoint.
ingestion_time Timestamp Auto-generated timestamp for when the data point was ingested into the Tilebox API.
value float64 A numeric measurement value.
sensor string A name of the sensor that generated the data point.
precise_time Timestamp A precise measurement time in nanosecond precision.
sensor_history Array[float64] The last few measurements of the sensor.

Once you’ve defined the schema and created a dataset, you can access it and create a collection to ingest data into.

Prepare data for ingestion

Ingestion is available in Python and Go.

Python

Every datapoint passed to collection.ingest must include time. Omit id and ingestion_time; Tilebox generates both fields during ingestion.

Record-oriented data

Use an iterable of mappings when you construct datapoints individually. Optional fields can be absent from individual records. None and common tabular missing values also leave optional fields unset.

Python
records = [
    {
        "time": "2025-03-28T11:44:23Z",
        "value": 45.16,
        "sensor": "A",
        "sensor_history": [-12.15, 13.45, -8.2, 16.5, 45.16],
    },
    {
        "time": "2025-03-28T11:45:19Z",
        "value": 273.15,
        "sensor": "B",
    },
]

datapoint_ids = collection.ingest(records)

Column-oriented data

Use a mapping of field names to equally sized sequences when your data is already organized by column.

Python
columns = {
    "time": [
        "2025-03-28T11:44:23Z",
        "2025-03-28T11:45:19Z",
    ],
    "value": [45.16, 273.15],
    "sensor": ["A", "B"],
}

collection.ingest(columns)

pandas DataFrame

Tilebox treats each DataFrame row as one datapoint and maps column names to dataset fields.

Python
import pandas as pd

data = pd.DataFrame({
    "time": [
        "2025-03-28T11:44:23Z",
        "2025-03-28T11:45:19Z",
    ],
    "value": [45.16, 273.15],
    "sensor": ["A", "B"],
})

collection.ingest(data)

xarray Dataset

Tilebox also accepts xarray.Dataset, the format returned when querying data.

Python
import numpy as np
import xarray as xr

data = xr.Dataset({
    "time": ("time", [
        "2025-03-28T11:46:13Z",
        "2025-03-28T11:46:54Z",
    ]),
    "value": ("time", [48.1, 290.12]),
    "sensor_history": (("time", "n_sensor_history"), [
        [13.45, -8.2, 16.5, 45.16, 48.1],
        [280.12, 273.15, 290.12, np.nan, np.nan],
    ]),
})

collection.ingest(data)

Go

Client.Datapoints.Ingest supports ingestion of data points in the form of a slice of protobuf messages.

Protobuf

Protobuf is Google’s language-neutral, platform-neutral, extensible mechanism for serializing structured data.

More details on protobuf can be found in the protobuf section.

In the example below, the v1.Modis type has been generated with tilebox dataset generate, as described in the protobuf section.

Go
datapoints := []*v1.Modis{
  v1.Modis_builder{
    Time:        timestamppb.New(time.Now()),
    GranuleName: proto.String("Granule 1"),
  }.Build(),
  v1.Modis_builder{
    Time:        timestamppb.New(time.Now().Add(-5 * time.Hour)),
    GranuleName: proto.String("Past Granule 2"),
  }.Build(),
}

ingestResponse, err := client.Datapoints.Ingest(ctx,
    collectionID,
    &datapoints
    false,
)

Copying or moving data

Since ingest takes query’s output as input, you can easily copy or move data from one collection to another.

Output plaintext
OtherMeasurements: [2025-03-28T11:44:23.000 UTC, 2025-03-28T11:46:54.000 UTC] (4 data points)

Automatic batching

Tilebox automatically batches the ingestion requests for you, so you don’t have to worry about the maximum request size.

Idempotency

Tilebox will auto-generate datapoint IDs based on the data of all its fields - except for the auto-generated ingestion_time, so ingesting the same data twice will result in the same ID being generated. By default, Tilebox will silently skip any data points that are duplicates of existing ones in a collection. This behavior is especially useful when implementing idempotent algorithms. That way, re-executions of certain ingestion tasks due to retries or other reasons will never result in duplicate data points.

You can instead also request an error to be raised if any of the generated datapoint IDs already exist. This can be done by setting the allow_existing parameter to False.

Output plaintext
ArgumentError: found existing datapoints with same id, refusing to ingest with "allow_existing=false"

Ingestion from common file formats

Through the usage of xarray and pandas you can also easily ingest existing datasets available in file formats, such as CSV, Parquet, Feather and more.

Check out the Ingestion from common file formats guide for examples of how to achieve this.

Assets

To ingest datapoints that reference files in external storage, see Reference assets in a dataset.

Geometries

Ingesting Geometries can traditionally be a bit tricky, especially when working with geometries that cross the antimeridian or cover a pole. Tilebox is designed to take away most of the friction involved in this, but it’s still recommended to follow the best practices for handling geometries.

Type to search…