# Ingesting Data

You need to have write permission on the collection to be able to ingest data.

Check out the examples below for common scenarios of ingesting data into a collection.

## Dataset schema

Tilebox Datasets are strongly typed. This means you can only ingest data that matches the schema of a dataset. The schema is defined during dataset creation time.

The examples on this page assume that you have access to a [Timeseries dataset](/docs/datasets/types/timeseries) that has the following schema:

**MyCustomDataset schema**

Check out the [Build a spatio-temporal catalog](/docs/guides/datasets/build-spatiotemporal-catalog) guide for an example of how to create such a dataset.

**MyCustomDataset schema**

| Field name       | Type            | Description                                                                                              |
| ---------------- | --------------- | -------------------------------------------------------------------------------------------------------- |
| `time`           | Timestamp       | Timestamp of the data point. Required by the [Timeseries dataset](/docs/datasets/types/timeseries) type. |
| `id`             | UUID            | Auto-generated UUID for each datapoint.                                                                  |
| `ingestion_time` | Timestamp       | Auto-generated timestamp for when the data point was ingested into the Tilebox API.                      |
| `value`          | float64         | A numeric measurement value.                                                                             |
| `sensor`         | string          | A name of the sensor that generated the data point.                                                      |
| `precise_time`   | Timestamp       | A precise measurement time in nanosecond precision.                                                      |
| `sensor_history` | Array\[float64] | The last few measurements of the sensor.                                                                 |

A full overview of available data types can be found in the [here](/docs/datasets/concepts/datasets#field-types).

Once you've defined the schema and created a dataset, you can access it and create a collection to ingest data into.

**Python**

```python title="Python"
from tilebox.datasets import Client

client = Client()
dataset = client.dataset("my_org.my_custom_dataset")
collection = dataset.get_or_create_collection("Measurements")
```

**Go**

```go title="Go"
package main

import (
	"context"
	"log"

	"github.com/tilebox/tilebox-go/datasets/v1"
)

func main() {
	ctx := context.Background()
	client := datasets.NewClient()

	dataset, err := client.Datasets.Get(ctx, "my_org.my_custom_dataset")
	if err != nil {
		log.Fatalf("Failed to get dataset: %v", err)
	}

	collection, err := client.Collections.GetOrCreate(ctx, dataset.ID, "Measurements")
	if err != nil {
		log.Fatalf("Failed to get collection: %v", err)
	}
}
```

## Prepare data for ingestion

Ingestion is available in Python and Go.

### Python

Every datapoint passed to [`collection.ingest`](/docs/api-reference/python/tilebox.datasets/Collection.ingest) must include `time`. Omit `id` and `ingestion_time`; Tilebox generates both fields during ingestion.

#### Record-oriented data

Use an iterable of mappings when you construct datapoints individually. Optional fields can be absent from individual records. `None` and common tabular missing values also leave optional fields unset.

```python title="Python"
records = [
    {
        "time": "2025-03-28T11:44:23Z",
        "value": 45.16,
        "sensor": "A",
        "sensor_history": [-12.15, 13.45, -8.2, 16.5, 45.16],
    },
    {
        "time": "2025-03-28T11:45:19Z",
        "value": 273.15,
        "sensor": "B",
    },
]

datapoint_ids = collection.ingest(records)
```

#### Column-oriented data

Use a mapping of field names to equally sized sequences when your data is already organized by column.

```python title="Python"
columns = {
    "time": [
        "2025-03-28T11:44:23Z",
        "2025-03-28T11:45:19Z",
    ],
    "value": [45.16, 273.15],
    "sensor": ["A", "B"],
}

collection.ingest(columns)
```

A mapping is always interpreted as column-oriented data. To ingest one record, wrap it in a list: `collection.ingest([record])`.

#### pandas DataFrame

Tilebox treats each DataFrame row as one datapoint and maps column names to dataset fields.

```python title="Python"
import pandas as pd

data = pd.DataFrame({
    "time": [
        "2025-03-28T11:44:23Z",
        "2025-03-28T11:45:19Z",
    ],
    "value": [45.16, 273.15],
    "sensor": ["A", "B"],
})

collection.ingest(data)
```

#### xarray Dataset

Tilebox also accepts [`xarray.Dataset`](/docs/sdks/python/xarray), the format returned when [querying data](/docs/datasets/query/querying-data).

```python title="Python"
import numpy as np
import xarray as xr

data = xr.Dataset({
    "time": ("time", [
        "2025-03-28T11:46:13Z",
        "2025-03-28T11:46:54Z",
    ]),
    "value": ("time", [48.1, 290.12]),
    "sensor_history": (("time", "n_sensor_history"), [
        [13.45, -8.2, 16.5, 45.16, 48.1],
        [280.12, 273.15, 290.12, np.nan, np.nan],
    ]),
})

collection.ingest(data)
```

Array fields use an extra xarray dimension, such as `n_sensor_history`. If array lengths differ, pad shorter values at the end with the fill value for that data type. Tilebox omits this trailing padding during ingestion.

### Go

[`Client.Datapoints.Ingest`](/docs/api-reference/go/datasets/Datapoints.Ingest) supports ingestion of data points in the form of a slice of protobuf messages.

#### Protobuf

Protobuf is Google's language-neutral, platform-neutral, extensible mechanism for serializing structured data.

More details on protobuf can be found in the [protobuf section](/docs/sdks/go/protobuf).

In the example below, the `v1.Modis` type has been generated with `tilebox dataset generate`, as described in the [protobuf section](/docs/sdks/go/protobuf).

```go title="Go"
datapoints := []*v1.Modis{
  v1.Modis_builder{
    Time:        timestamppb.New(time.Now()),
    GranuleName: proto.String("Granule 1"),
  }.Build(),
  v1.Modis_builder{
    Time:        timestamppb.New(time.Now().Add(-5 * time.Hour)),
    GranuleName: proto.String("Past Granule 2"),
  }.Build(),
}

ingestResponse, err := client.Datapoints.Ingest(ctx,
    collectionID,
    &datapoints
    false,
)
```

## Copying or moving data

Since `ingest` takes `query`'s output as input, you can easily copy or move data from one collection to another.

Copying data like this also works across datasets in case the dataset schemas are compatible.

**Python**

```python title="Python"
src_collection = dataset.collection("Measurements")
data_to_copy = src_collection.query(temporal_extent=("2025-03-28", "2025-03-29"))

dest_collection = dataset.collection("OtherMeasurements")
dest_collection.ingest(data_to_copy)  # copy the data to the other collection

# To verify it now contains 4 datapoints (2 we ingested already, and 2 we copied just now)
print(dest_collection.info())
```

**Go**

```go title="Go"
dataset, err := client.Datasets.Get(ctx, "my_org.my_custom_dataset")
if err != nil {
  log.Fatalf("Failed to get dataset: %v", err)
}

srcCollection, err := client.Collections.GetOrCreate(ctx, dataset.ID, "Measurements")
if err != nil {
  log.Fatalf("Failed to get collection: %v", err)
}

startDate := time.Date(2025, time.March, 28, 0, 0, 0, 0, time.UTC)
endDate := time.Date(2025, time.March, 29, 0, 0, 0, 0, time.UTC)

var dataToCopy []*v1.MyCustomDataset
err = client.Datapoints.QueryInto(ctx,
  dataset.ID,
  &dataToCopy,
  datasets.WithCollectionIDs(srcCollection.ID),
  datasets.WithTemporalExtent(query.NewTimeInterval(startDate, endDate)),
)
if err != nil {
  log.Fatalf("Failed to query datapoints: %v", err)
}

destCollection, err := client.Collections.GetOrCreate(ctx, dataset.ID, "OtherMeasurements")
if err != nil {
  log.Fatalf("Failed to get collection: %v", err)
}

// copy the data to the other collection
_, err = client.Datapoints.Ingest(ctx, destCollection.ID, &dataToCopy, false)
if err != nil {
  log.Fatalf("Failed to ingest datapoints: %v", err)
}

// To verify it now contains 4 datapoints (2 we ingested already, and 2 we copied just now)
updatedDestCollection, err := client.Collections.Get(ctx, dataset.ID, "OtherMeasurements")
if err != nil {
  log.Fatalf("Failed to get collection: %v", err)
}
slog.Info("Updated collection", slog.String("collection", updatedDestCollection.String()))
```

**Output**

```plaintext title="Output"
OtherMeasurements: [2025-03-28T11:44:23.000 UTC, 2025-03-28T11:46:54.000 UTC] (4 data points)
```

## Automatic batching

Tilebox automatically batches the ingestion requests for you, so you don't have to worry about the maximum request size.

## Idempotency

Tilebox will auto-generate datapoint IDs based on the data of all its fields - except for the auto-generated
`ingestion_time`, so ingesting the same data twice will result in the same ID being generated. By default, Tilebox
will silently skip any data points that are duplicates of existing ones in a collection. This behavior is especially
useful when implementing idempotent algorithms. That way, re-executions of certain ingestion tasks due to retries
or other reasons will never result in duplicate data points.

You can instead also request an error to be raised if any of the generated datapoint IDs already exist.
This can be done by setting the `allow_existing` parameter to `False`.

**Python**

```python title="Python"
data = pd.DataFrame({
    "time": [
      "2025-03-28T11:45:19Z",
    ],
    "value": [45.16],
    "sensor": ["A"],
    "precise_time": [
      "2025-03-28T11:44:23.345761444Z",
    ],
    "sensor_history": [
      [-12.15, 13.45, -8.2, 16.5, 45.16],
    ],
})

# we already ingested the same data point previously
collection.ingest(data, allow_existing=False)

# we can still ingest it, by setting allow_existing=True
# but the total number of datapoints will still be the same
# as before in that case, since it already exists and therefore
# will be skipped
collection.ingest(data, allow_existing=True)  # no-op
```

**Go**

```go title="Go"
datapoints := []*v1.MyCustomDataset{
  v1.MyCustomDataset_builder{
    Time:          timestamppb.New(time.Date(2025, time.March, 28, 11, 45, 19, 0, time.UTC)),
    Value:         proto.Float64(45.16),
    Sensor:        proto.String("A"),
    PreciseTime:   timestamppb.New(time.Date(2025, time.March, 28, 11, 44, 23, 345761444, time.UTC)),
    SensorHistory: []float64{-12.15, 13.45, -8.2, 16.5, 45.16},
  }.Build(),
}

// we already ingested the same data point previously
_, err = client.Datapoints.Ingest(ctx, collection.ID, &datapoints, false)
if err != nil {
  log.Fatalf("Failed to ingest datapoints: %v", err)
}

// we can still ingest it, by setting allowExisting to true
// but the total number of datapoints will still be the same
// as before in that case, since it already exists and therefore
// will be skipped
_, err = client.Datapoints.Ingest(ctx, collection.ID, &datapoints, true) // no-op
if err != nil {
  log.Fatalf("Failed to ingest datapoints: %v", err)
}
```

**Output**

```plaintext title="Output"
ArgumentError: found existing datapoints with same id, refusing to ingest with "allow_existing=false"
```

## Ingestion from common file formats

Through the usage of `xarray` and `pandas` you can also easily ingest existing datasets available in file
formats, such as CSV, [Parquet](https://parquet.apache.org/), [Feather](https://arrow.apache.org/docs/python/feather.html) and more.

Check out the [Ingestion from common file formats](/docs/guides/datasets/ingest-format) guide for examples of how to achieve this.

## Assets

To ingest datapoints that reference files in external storage, see [Reference assets in a dataset](/docs/datasets/assets-and-storage/reference-assets).

## Geometries

Ingesting Geometries can traditionally be a bit tricky, especially when working with geometries that cross the antimeridian or cover a pole.
Tilebox is designed to take away most of the friction involved in this, but it's still recommended to follow the [best practices for handling geometries](/docs/datasets/geometries).
