<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Tilebox Articles</title><link>https://tilebox.com/dispatch/articles</link><description>Technical guides, explanations, and perspectives on satellite data and geospatial workflows.</description><language>en</language><image><url>https://tilebox.com/rss-icon.png</url><title>Tilebox Articles</title><link>https://tilebox.com/dispatch/articles</link><width>144</width><height>144</height></image><atom:link href="https://tilebox.com/dispatch/articles/rss.xml" rel="self" type="application/rss+xml"/><item><title>GeoAI Can&apos;t Scale On Models Alone: It Needs An Operating Layer</title><link>https://tilebox.com/dispatch/articles/geoai-needs-an-operating-layer</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/geoai-needs-an-operating-layer</guid><description>GeoAI needs an operating layer that solves existing engineering bottlenecks to deliver on its real potential at scale</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/cfb1ee42ff5d82c7.webp&quot; alt=&quot;GeoAI Can&apos;t Scale On Models Alone: It Needs An Operating Layer&quot;&gt;&lt;/p&gt;&lt;h2 id=&quot;geoai-needs-to-get-grounded&quot;&gt;GeoAI needs to get grounded&lt;/h2&gt;
&lt;p&gt;GeoAI is moving fast because the opportunity is real. Satellite imagery, SAR, hyperspectral data, foundation models, and AI agents are making it possible to ask operational questions about the physical world and get answers in minutes.&lt;/p&gt;
&lt;p&gt;The market is no longer theoretical. Planet is turning daily Earth observation into monitoring contracts. ICEYE has become one of the clearest signals of demand for persistent, all-weather intelligence. BlackSky’s Spectra AI platform points to the shift from imagery to real-time monitoring. Pixxel is adding hyperspectral data to the stack. Kayrros turns satellite observations into climate and energy intelligence. Xoople is framing the next phase directly: mapping the Earth for AI. Google, Microsoft, and IBM are pushing geospatial foundation models closer to the tools analysts already use.&lt;/p&gt;
&lt;p&gt;While there is ample opportunity, the limitation is that more models, more satellites, and more data do not automatically create grounded workflows.&lt;/p&gt;
&lt;h2 id=&quot;the-stack-can-infer-it-still-needs-to-verify&quot;&gt;The stack can infer. It still needs to verify.&lt;/h2&gt;
&lt;p&gt;The leading GeoAI companies make different parts of the stack stronger: better imagery, richer sensors, faster revisit, stronger analytics, easier model access, and more verticalized signals.&lt;/p&gt;
&lt;p&gt;But the hard part is no longer proving that GeoAI can produce an answer. The hard part is proving that the answer can be trusted, reproduced, and operationalized.&lt;/p&gt;
&lt;p&gt;An AI agent can sound right without being right.&lt;/p&gt;
&lt;p&gt;Geospatial data is numeric, spatial, and consequential. A wrong field name, stale dataset, bad geometry predicate, or unverified flood extent can affect a risk model, infrastructure plan, supply-chain forecast, or public-sector response.&lt;/p&gt;
&lt;p&gt;Three major limitations show up across the market.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The context is fragmented.&lt;/strong&gt; Every provider has its own catalog, schema, access pattern, license, format, and cloud footprint. A GeoAI team can have excellent data and still spend too much time translating between systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The workflows are brittle.&lt;/strong&gt; Agents can generate plausible-looking code that references fields that do not exist, assumes the wrong time column, or breaks when moved to a new region. Foundation models can perform well on familiar data and degrade when the geography, sensor, or local context changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The outputs are hard to verify.&lt;/strong&gt; A model output is not an operational record. For high-stakes use cases, teams need to know what ran, which data it used, what parameters changed, and what the system produced.&lt;/p&gt;
&lt;p&gt;The market is scaling its promises faster than it is scaling its ability to ground them.&lt;/p&gt;
&lt;h2 id=&quot;geoai-needs-an-operating-layer&quot;&gt;GeoAI needs an operating layer&lt;/h2&gt;
&lt;p&gt;The next phase of GeoAI will not be won by the model alone. It will be won by teams that turn AI outputs into governed, repeatable workflows.&lt;/p&gt;
&lt;p&gt;That requires an operating layer with guardrails for agents. And that’s where the Tilebox agentic framework fits in. It gives coding agents the operational context and engineering tools they need to work with geospatial data safely and repeatably.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Live context.&lt;/strong&gt; Agents inspect real dataset schemas, collections, query options, workflow state, jobs, logs, and spans before acting. They do not have to guess whether the field is &lt;code&gt;cloudcover&lt;/code&gt; or &lt;code&gt;cloud_cover&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Deterministic tools.&lt;/strong&gt; Agents use the Tilebox CLI to run real commands with machine-readable inputs and outputs. They can query data, submit jobs, deploy workflows, inspect failures, and iterate from the terminal.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent skills.&lt;/strong&gt; Tilebox skills teach agents how to manage datasets, monitor jobs, write workflows, release code, and inspect automations. The agent gets the operating pattern once and applies it across projects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distributed compute.&lt;/strong&gt; Tilebox lets agents run geospatial workflows closer to the data, reducing data movement and one-off infrastructure work as teams expand into new regions or products.&lt;/p&gt;
&lt;p&gt;GeoAI will transform global, regional, and local intelligence. But it needs more than signals and models. It needs agents that can act through governed systems, inspect what happened, and reproduce the result.&lt;/p&gt;
&lt;p&gt;Tilebox is how GeoAI gets grounded.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/cfb1ee42ff5d82c7.webp" medium="image"><media:title>GeoAI Can&apos;t Scale On Models Alone: It Needs An Operating Layer</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/cfb1ee42ff5d82c7.webp"/><category>article</category><pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate><dc:creator>Megan Van Patten</dc:creator></item><item><title>Climate Tech&apos;s Most Unlikely Ally</title><link>https://tilebox.com/dispatch/articles/climate-tech</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/climate-tech</guid><description>Tilebox-skilled agents offer efficient, verified, reproducible geospatial workflows to diversify and scale timely climate analytics.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/4a51e66d95039a0c.webp&quot; alt=&quot;Climate Tech&apos;s Most Unlikely Ally&quot;&gt;&lt;/p&gt;&lt;p&gt;Tilebox-skilled agents offer efficient, verified, reproducible geospatial workflows to diversify and scale climate analytics for timely action.&lt;/p&gt;
&lt;h2 id=&quot;the-climate-tech-conundrum&quot;&gt;The climate tech conundrum&lt;/h2&gt;
&lt;p&gt;Climate risk analytics, extreme-weather forecasting, water and nature solutions was the fastest-growing segment in climate tech in 2025, its share of funding climbing from 5.0% to 8.2% in a single year on a projected $5.5 billion (Net Zero Insights, &lt;a href=&quot;https://stateofclimatetech.com/&quot;&gt;State of Climate Tech 2025&lt;/a&gt;). Insurers, asset managers, and corporates filing TCFD and IFRS S2 disclosures are all hungry for asset-level risk insight.&lt;/p&gt;
&lt;p&gt;And yet, while US climate-tech investment reached $29 billion in 2025, a record 52% of venture-backed climate-tech companies cut their net burn year-over-year just to stay resilient (Silicon Valley Bank, &lt;a href=&quot;https://www.svb.com/trends-insights/reports/future-of-climate-tech/&quot;&gt;Future of Climate Tech 2026&lt;/a&gt;). The promise is enormous, but the economics are unforgiving. The difference between the two is almost always engineering overhead.&lt;/p&gt;
&lt;h2 id=&quot;fragile-scaffolding&quot;&gt;Fragile scaffolding&lt;/h2&gt;
&lt;p&gt;The geospatial market sorts into three tiers. At the bottom sit the EO satellite operators (Planet, ICEYE, Vantor) plus public archives from NASA and ESA, supplying raw imagery. In the middle, hyperscalers like Google, Microsoft, and AWS store the petabytes and ship powerful geospatial foundation models. At the top is the intelligence layer: the companies that fuse EO data with weather records, economic data, and proprietary models to deliver something a risk manager can actually act on.&lt;/p&gt;
&lt;p&gt;That top layer is where most of the value, and complexity, concentrates. It’s also where the engineering burden is heaviest, because intelligence-layer companies inherit every messy seam between the layers below them: incompatible formats, authentication to a dozen archives, and enormous files that cost real money every time they cross a cloud region.&lt;/p&gt;
&lt;p&gt;The egress math is not hypothetical — at standard cloud pricing of roughly $0.09/GB, moving a single large embedding dataset out of one provider has run teams tens of thousands of dollars in transfer fees alone, and bespoke formats compound that tax across every downstream product. Not to mention, value-killing latency which adds to the challenge of these climate orgs making enough to keep the lights on.&lt;/p&gt;
&lt;h2 id=&quot;where-foundation-models-stop-and-local-intelligence-begins&quot;&gt;Where foundation models stop and local intelligence begins&lt;/h2&gt;
&lt;p&gt;Geospatial foundation models from Google, IBM (Prithvi), Microsoft (Aurora), and Clay have been a genuine leap, compressing raw imagery into representations downstream apps can use at a fraction of the previous compute cost. But even as these models get richer, the value in climate intelligence is rarely global. It’s regional and specific: this watershed, this county’s flood exposure, this cooperative’s fields.&lt;/p&gt;
&lt;p&gt;On independent benchmarks, a model that performs competitively on familiar data can collapse to under 20% accuracy when moved to a new region (&lt;a href=&quot;https://arxiv.org/html/2412.04204v1&quot;&gt;PANGAEA&lt;/a&gt;). &lt;a href=&quot;https://tilebox.com/dispatch/articles/parametric-insurance&quot;&gt;Parametric insurance&lt;/a&gt; runs into the same wall from the other direction: the events it pays out on are inherently local, so a model that’s only right on average is wrong exactly where a claim is filed.&lt;/p&gt;
&lt;p&gt;Delivering actionable value takes local sensor data, regional context, validation against ground truth, and constant re-tuning, without a custom build each time. The hard part isn’t running the model. It’s feeding it the right local data continuously, at the granularity a region demands. And that is an orchestration problem.&lt;/p&gt;
&lt;h2 id=&quot;fixing-whats-broken&quot;&gt;Fixing what’s broken&lt;/h2&gt;
&lt;p&gt;Three failure modes recur across climate data companies, and none of them are about the science. These are the common challenges across geospatial development that Tilebox was founded to solve.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Pipeline sprawl.&lt;/strong&gt; Every new product pulls from different sources in different formats, so each becomes a fragile custom pipeline to maintain, and the cost of a product line is dominated not by the model but by the connectors feeding it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The geospatial data tax.&lt;/strong&gt; Imagery files are huge, moving them between regions burns egress fees, and wiring up connectors takes weeks of engineering before a single insight ships — overhead that, for a small team, is the difference between profitability and another bridge round.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The last-mile workflow gap.&lt;/strong&gt; Since finance runs on tables and feeds rather than maps, insight that can’t drop into the format and cadence a client already works in doesn’t matter how good the underlying data is.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To fix these problems, you need both interoperability and speed. That’s why we released Tilebox for agentic engineering, and built it to support any data source, environment, algorithm, and coding agent. More importantly, when using Tilebox, agents work more efficiently.&lt;/p&gt;
&lt;h2 id=&quot;agents-that-run-the-pipeline-not-just-the-model&quot;&gt;Agents that run the pipeline, not just the model&lt;/h2&gt;
&lt;p&gt;The difference between an AI answer and an operational workflow is whether you can inspect, reproduce, and trust the result. Agents on Tilebox use the same operational context human developers do, through three primitives:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;MCP server&lt;/strong&gt; exposes live dataset schemas, query tools, and job status. The agent reads the real schema (exact fields like &lt;code&gt;cloud_cover&lt;/code&gt; , &lt;code&gt;precise_time&lt;/code&gt; ) instead of guessing from stale docs, so its code runs first time rather than failing on an incorrect field.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CLI&lt;/strong&gt; machine-readable descriptions of every command and return. The agent runs real, governed actions (discover data, trigger a workflow, inspect execution) then parses the result and repeats the loop. No human stitching steps together; the run is reviewable from the terminal.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agent skills&lt;/strong&gt; a single-line install drops them into the universal skills directory your coding agent already reads (Claude Code, Cursor, Codex). No special prompt prefix: install once, describe the task in plain language.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The result: agents act through real commands against governed data, leaving a record of what ran, where, on which data, and what it produced.&lt;/p&gt;
&lt;p&gt;Grounded that way, a climate analytics company’s data operations can run largely autonomously: agents query open (or custom) datasets, deploy risk and monitoring pipelines, adjust model thresholds, and answer “are the GPU tasks stuck on the cloud cluster?” from live task state instead of a manual log dig.&lt;/p&gt;
&lt;p&gt;And because Tilebox’s architecture is inherently built for distributed computing, agents catalog and transform data without moving it. For a marketplace model, that autonomy extends to the catalog itself: agents can onboard and validate new datasets far faster than a manual review queue allows.&lt;/p&gt;
&lt;h2 id=&quot;start-building&quot;&gt;Start building&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://datacenterbuildout.com/?utm_source=blog&amp;amp;utm_medium=climate&amp;amp;utm_campaign=agent_launch&quot;&gt;Here’s a change detection tracker&lt;/a&gt; we built with an agent. Watch the demo video to see how to build the workflow underneath then run it yourself or fork the &lt;a href=&quot;https://github.com/tilebox/datacenters?utm_source=blog&amp;amp;utm_medium=climate&amp;amp;utm_campaign=agent_launch&quot;&gt;source code&lt;/a&gt;.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/4a51e66d95039a0c.webp" medium="image"><media:title>Climate Tech&apos;s Most Unlikely Ally</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/4a51e66d95039a0c.webp"/><category>article</category><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><dc:creator>Megan Van Patten</dc:creator></item><item><title>Parametric Insurance That Scales</title><link>https://tilebox.com/dispatch/articles/parametric-insurance</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/parametric-insurance</guid><description>Parametric insurance promises a faster, fairer model for managing climate risk. Tilebox provides the agent-ready infrastructure that finally lets it scale.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/54dd40347a088a96.webp&quot; alt=&quot;Parametric Insurance That Scales&quot;&gt;&lt;/p&gt;&lt;p&gt;The parametric insurance model promises instant payouts the moment a hurricane crosses a windspeed threshold or a region misses its rainfall benchmark. This new era of insurance is poised to provide greater clarity for customers, but the back-end operations haven’t been so straight forward — or financially sufficient.&lt;/p&gt;
&lt;h2 id=&quot;bottlenecks-and-bottom-lines&quot;&gt;Bottlenecks and bottom lines&lt;/h2&gt;
&lt;p&gt;Investments in machine learning have sharpened the algorithms and replaced latent claims processing. However, a processing bottleneck persists, just with a different input: multi-source geospatial data. A mess of customized, fragile pipelines stitching together several packages that never work as seamlessly as stakeholders demand they do.&lt;/p&gt;
&lt;p&gt;Tileboxchanges the economics of that workload by letting AI agents run the geospatial pipeline themselves, using an ultra-efficient and agnostic framework built for humans and agents.&lt;/p&gt;
&lt;h2 id=&quot;the-geospatial-data-tax&quot;&gt;The “geospatial data tax”&lt;/h2&gt;
&lt;p&gt;The engineering overhead required to monitor global assets is often referred to as a &lt;em&gt;geospatial data tax&lt;/em&gt;. Satellite imagery files are enormous. Moving them between cloud regions burns through egress fees. Building custom connectors to public archives and private constellations requires weeks, if not months, of complex cloud and data engineering. From drought coverage for Kenyan farmers to wildfires for California vineyards, every new product means another custom build. Another custom build requires more manual maintenance.&lt;/p&gt;
&lt;p&gt;The ability to scale parametric algorithms for optimal proficiency is fumbled by the fragmentation of each pipeline. Tilebox removes these burdens and offers a simplified entry point to geospatial data processing for new and evolving markets.&lt;/p&gt;
&lt;h2 id=&quot;agent-see-agent-do&quot;&gt;Agent see, agent do&lt;/h2&gt;
&lt;p&gt;With Tilebox, AI agents access the same operational context developers use: live documentation, authenticated access to Tilebox resources, and deterministic tools for running real actions. This means an insurance company’s data operations can become largely autonomous. Agents automate new event data ingestion, deploy new risk pipelines, and adjust thresholds without human DevOps in the loop.&lt;/p&gt;
&lt;h2 id=&quot;the-roadmap-to-real-value&quot;&gt;The roadmap to real value&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Efficiency.&lt;/strong&gt; Parametric insurers don’t need to build entire space-data engineering departments. Agents handle ingestion, cataloging, and pipeline maintenance on Tilebox’s process-where-it-sits architecture, avoiding costly data movement, vendor dependencies, and wasted engineering resources.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Market growth.&lt;/strong&gt; When launching a new parametric product no longer requires six months of engineering work, insurers can serve markets that were previously uneconomic: smallholder farmers, SMB flood cover, hyper-local windstorm policies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Faster payouts.&lt;/strong&gt; Threshold breaches calculated and verified during the satellite pass itself can automate near-instant settlement. Talk with our team about Tilebox On Orbit for edge computing payloads.&lt;/p&gt;
&lt;p&gt;As the majority of the EO sector races to hammer out siloed GeoAI use cases, agents running on Tilebox can deliver repeatable workflows with deterministic results across markets.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/54dd40347a088a96.webp" medium="image"><media:title>Parametric Insurance That Scales</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/54dd40347a088a96.webp"/><category>article</category><pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate><dc:creator>Megan Van Patten</dc:creator></item><item><title>Are You A Platform Integrator Or An EO Developer?</title><link>https://tilebox.com/dispatch/articles/hidden-costs-of-fragmentation</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/hidden-costs-of-fragmentation</guid><description>Eliminate the undifferentiated work and focus on developing and scaling your geospatial product.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/88b479589bf3d84c.webp&quot; alt=&quot;Are You A Platform Integrator Or An EO Developer?&quot;&gt;&lt;/p&gt;&lt;p&gt;Every geospatial data pipeline needs the same foundational capabilities:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A way to query metadata across multiple providers&lt;/li&gt;
&lt;li&gt;A workflow engine that handles parallel execution and resumes from failure&lt;/li&gt;
&lt;li&gt;Triggers that fire when new data lands&lt;/li&gt;
&lt;li&gt;Code that’s easy enough to write so you can iterate quickly&lt;/li&gt;
&lt;li&gt;As hyperspectral data adoption rises, a way to retrieve only the bands you actually need&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;None of these requirements are exotic. The problem is that no single tool has historically delivered all of them; teams build their stack by stitching together a STAC client, a PostGIS database, an Airflow or Temporal instance, custom storage retrieval scripts per provider, and homegrown deduplication logic on top.&lt;/p&gt;
&lt;p&gt;That’s five tools, five sets of credentials, five places where things can break, and five ongoing maintenance burdens. Before a team writes a single line of the algorithm that actually differentiates their product, they’ve signed up to be a platform integrator.&lt;/p&gt;
&lt;h2 id=&quot;the-hidden-costs-of-fragmentation&quot;&gt;The hidden costs of fragmentation&lt;/h2&gt;
&lt;p&gt;The compounding cost of this fragmented stack shows up in three places.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Engineering time goes to integration work that no customer ever sees.&lt;/li&gt;
&lt;li&gt;New team members spend weeks learning a custom stack instead of a standard one.&lt;/li&gt;
&lt;li&gt;Every provider API change or breaking update means rewriting integration code your team never budgeted time for.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Every hour spent gluing tools together is an hour not spent on the algorithm, the user experience, or the customer. Teams under deadline pressure inevitably make trade-offs against infrastructure quality, and over time, those trade-offs compound into slower iteration, harder onboarding, and &lt;strong&gt;your competitive position becomes defined by what the team can maintain&lt;/strong&gt; rather than what they could be building.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/31b207b7b07eb913.webp&quot; alt=&quot;Engineering time split between infrastructure and product development, compared in two donut charts&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;a-different-architecture&quot;&gt;A different architecture&lt;/h2&gt;
&lt;p&gt;Tilebox combines those capabilities into a single framework. We created a developer experience that makes building faster with an execution architecture that makes running pipelines more cost-efficient and more reliable.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Spatio-temporal metadata querying across open data and proprietary datasets through one API&lt;/li&gt;
&lt;li&gt;Storage clients that retrieve files directly from provider infrastructure with partial product downloads built in&lt;/li&gt;
&lt;li&gt;A workflow engine with automatic parallelization, retries, and idempotency&lt;/li&gt;
&lt;li&gt;Storage event and cron-based automations for near-real-time processing&lt;/li&gt;
&lt;li&gt;Multi-language SDKs (Python, Go) and an MCP server for AI-assisted development&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each of these capabilities exists somewhere else. None of them exist together, built specifically for EO, in a single integrated framework - until we launched Tilebox.&lt;/p&gt;
&lt;h2 id=&quot;focus-on-the-work-that-matters&quot;&gt;Focus on the work that matters&lt;/h2&gt;
&lt;p&gt;When the infrastructure stack collapses from five tools to one, the work that disappears is the undifferentiated work. What’s left is the work that actually matters: the algorithms, the analytics, the products.&lt;/p&gt;
&lt;p&gt;Teams using Tilebox work faster because the data infrastructure already exists. They’re building a product, not a product &lt;em&gt;and&lt;/em&gt; the infrastructure underneath it. They’re spending their time on the things customers actually pay for, and deploying production-grade EO products in months rather than years.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/88b479589bf3d84c.webp" medium="image"><media:title>Are You A Platform Integrator Or An EO Developer?</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/88b479589bf3d84c.webp"/><category>article</category><pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate><dc:creator>Megan Van Patten</dc:creator></item><item><title>How to Simplify Multi-Source EO Data Pipelines</title><link>https://tilebox.com/dispatch/articles/how-to-simplify-multi-source-eo-data-pipelines</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/how-to-simplify-multi-source-eo-data-pipelines</guid><description>Skip weeks of integration work. With Tilebox, a new data source is just a new dataset name and one additional line of code.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/678ece70a19df7b8.webp&quot; alt=&quot;How to Simplify Multi-Source EO Data Pipelines&quot;&gt;&lt;/p&gt;
&lt;p&gt;No single satellite sensor can do everything. Optical cameras go blind in clouds or at night. Some sensors can detect crop stress but can’t explain why. Others see through clouds but miss important visual detail. As the idiom goes, more sensors are better than one.&lt;/p&gt;
&lt;p&gt;Combining data from multiple satellites is technically hard and expensive. Teams spend months on custom pipeline “plumbing” just to handle different data formats, failed downloads, and scaling issues, leaving little time for real science. Many end up cutting scope entirely, limiting themselves to a single data source to avoid the headache.&lt;/p&gt;
&lt;h2 id=&quot;evolving-the-geospatial-workflow&quot;&gt;Evolving the Geospatial Workflow&lt;/h2&gt;
&lt;p&gt;Given the exponential growth in satellite data and relatively low ROI for downstream geospatial applications, it’s important to point out that the bottleneck is not sales. It’s the pipeline. Between the total acquisition cost of data, the engineering resources it takes to build and maintain pipelines, plus growing competition - its critical to focus on differentiation and offering the best available product using the best available data, no matter the source(s).&lt;/p&gt;
&lt;p&gt;We’re working to evolve the geospatial sector by solving for traditional obstacles and drastically improving data efficiency.&lt;/p&gt;
&lt;h2 id=&quot;here-are-a-few-of-the-things-were-changing&quot;&gt;Here are a few of the things we’re changing:&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One framework for any environment(s).&lt;/strong&gt; If you need to process Sentinel-2 data from an Open Telekom Cloud, Landsat data from AWS US-West-2, and proprietary data in a Google Cloud for a multi-source pipeline, Tilebox enables you to execute the same workflow in each environment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Parallel processing at scale.&lt;/strong&gt; One job can automatically split into thousands of smaller tasks, no manual batching required.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Automatic error recovery.&lt;/strong&gt; If a download fails, Tilebox retries it. A 10,000-scene job that breaks at scene 8,743 resumes from there, not from the beginning.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Runs on your infrastructure.&lt;/strong&gt; You control the compute and cost; Tilebox handles the coordination.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Adding a new data source used to mean weeks of integration work. With Tilebox, it’s just a new dataset name and one additional line of code.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;client.dataset(&quot;&amp;lt;slug&amp;gt;&quot;)&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;For EO teams, that’s the difference between multi-source data being a real advantage, or just a project that never gets built.&lt;/p&gt;
&lt;p&gt;Check out this video for a sample workflow.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://tilebox.com/images/publication/9c7ef06b202eb189.webp&quot; alt=&quot;&quot;&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=uWyKuRxW_EE&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/figure&gt; </content:encoded><media:content url="https://tilebox.com/images/publication/678ece70a19df7b8.webp" medium="image"><media:title>How to Simplify Multi-Source EO Data Pipelines</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/678ece70a19df7b8.webp"/><category>article</category><pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate><dc:creator>Megan Van Patten</dc:creator></item><item><title>3 Ways to Start Building With Tilebox: From Query to Code</title><link>https://tilebox.com/dispatch/articles/3-ways-to-start-building-with-tilebox-from-query-to-code</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/3-ways-to-start-building-with-tilebox-from-query-to-code</guid><description>Whether you are exploring geospatial data for the first time, or a veteran engineer, here are three ways Tilebox makes it easier to get from query to code.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/a9769484884c233b.webp&quot; alt=&quot;3 Ways to Start Building With Tilebox: From Query to Code&quot;&gt;&lt;/p&gt;&lt;p&gt;Whether you are exploring geospatial data for the first time, or a veteran engineer, here are three ways Tilebox makes it easier to get from query to code.&lt;/p&gt;
&lt;h2 id=&quot;1-export-as-code-in-the-console&quot;&gt;1. Export as Code in the Console&lt;/h2&gt;
&lt;p&gt;The &lt;a href=&quot;https://docs.tilebox.com/console&quot;&gt;Tilebox Console&lt;/a&gt; is a visual dataset explorer. Browse available datasets, draw your area of interest, set a time window, apply quality filters, and then hit &lt;strong&gt;Export as Code&lt;/strong&gt;. You get a ready-to-run Python script, pre-filled with the correct dataset slug, collection name, and storage client for that provider. Copy, paste, run.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Best for: first-time users, analysts who want to prototype quickly, or anyone exploring a dataset they haven’t worked with before.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/581072fd997f0448.webp&quot; alt=&quot;Tilebox Console Export as Code feature&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;2-write-it-directly-with-the-sdk&quot;&gt;2. Write It Directly with the SDK&lt;/h2&gt;
&lt;p&gt;Once you’re comfortable with &lt;a href=&quot;https://docs.tilebox.com/&quot;&gt;Tilebox’s patterns&lt;/a&gt;, you can write queries directly. Install &lt;code&gt;tilebox-datasets&lt;/code&gt; and &lt;code&gt;tilebox-storage&lt;/code&gt;, initialize a client with your API key, and you’re querying scene metadata with a spatio-temporal filter in a few lines. The same query structure works across every dataset in your catalog (Sentinel, Landsat, Umbra SAR, Wyvern). Write once, deploy anywhere.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Best for: developers building pipelines, automating workflows, or integrating Tilebox into existing codebases.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;3-ask-an-agent&quot;&gt;3. Ask an Agent&lt;/h2&gt;
&lt;p&gt;After connecting the &lt;a href=&quot;https://docs.tilebox.com/agents-and-ai-tools/tilebox-mcp&quot;&gt;Tilebox MCP&lt;/a&gt; directly to your AI coding assistant, like Claude or Cursor, the agent can query your live catalog in real time. It knows your dataset slugs, collection names, and schemas before it writes a single line, so you can prompt it in plain English: &lt;em&gt;“Write me a script to pull all Sentinel-2 scenes over Berlin with less than 10% cloud cover for Q1 2025.”&lt;/em&gt; The agent handles the lookup problem entirely.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Best for: developers working in AI-assisted IDEs, and anyone who wants to skip from intent to working code in one step.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Whichever path fits your workflow, the underlying result is the same: less time on data access infrastructure, more time on the analysis that actually matters.&lt;/p&gt;
&lt;p&gt;Start prototyping or validate your current workflows with Tilebox Labs | $0/mo - no credit card required. &lt;a href=&quot;https://console.tilebox.com/sign-up&quot;&gt;Sign up here.&lt;/a&gt;&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/a9769484884c233b.webp" medium="image"><media:title>3 Ways to Start Building With Tilebox: From Query to Code</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/a9769484884c233b.webp"/><category>article</category><pubDate>Wed, 22 Apr 2026 00:00:00 GMT</pubDate><dc:creator>Megan Van Patten</dc:creator></item><item><title>How to Query Sentinel-2 Data in Python (No Setup Required)</title><link>https://tilebox.com/dispatch/articles/how-to-query-sentinel-2-data-in-python</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/how-to-query-sentinel-2-data-in-python</guid><description>Query Sentinel-2 scenes by location, time, and cloud cover in five lines of Python. One package, no ESA account or API credentials. Full script included.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/3583041da4f66b08.webp&quot; alt=&quot;How to Query Sentinel-2 Data in Python (No Setup Required)&quot;&gt;&lt;/p&gt;&lt;p&gt;In just five lines of Python, you can query every Sentinel-2 scene over any location on Earth – filtered by time, geography, and cloud cover — with the Tilebox Open Data catalog.&lt;/p&gt;
&lt;p&gt;No ESA account, no Copernicus API credentials, no OAuth dance required. Install one package, point it at the Tilebox Open Data catalog, and get an xarray Dataset back.&lt;/p&gt;
&lt;p&gt;The output is an interactive folium map showing Sentinel-2 scene footprints over your area of interest, colored by cloud cover percentage. &lt;em&gt;Open in Google Colab&lt;/em&gt; &lt;a href=&quot;https://colab.research.google.com/drive/1_6aBoXl8RhPRnz3hi33lIrhZkBH5bTx_#scrollTo=RaDy6OFLBz9o&quot;&gt;&lt;em&gt;here&lt;/em&gt;&lt;/a&gt;&lt;em&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/d9cedaf61f97f511.webp&quot; alt=&quot;Sentinel-2 scene footprints on a map&quot;&gt;&lt;img src=&quot;https://tilebox.com/images/publication/36d533d2f4c7b9cd.webp&quot; alt=&quot;Sentinel-2 scene thumbnail&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;install-tilebox&quot;&gt;Install Tilebox&lt;/h2&gt;
&lt;pre&gt;&lt;code&gt;uv add tilebox    # or: pip install tilebox&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That’s the only dependency you need for querying. We’ll add &lt;code&gt;folium&lt;/code&gt; and &lt;code&gt;matplotlib&lt;/code&gt; later for visualization.&lt;/p&gt;
&lt;h2 id=&quot;query-sentinel-2-scenes-over-a-location&quot;&gt;Query Sentinel-2 scenes over a location&lt;/h2&gt;
&lt;p&gt;Pick a location and a time range. This example queries all Sentinel-2 Level 2A scenes over Paris during September 2025.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;from shapely import Polygon
from tilebox.datasets import Client

# Polygon around Paris
paris = Polygon([
    (2.245958, 48.8311913), (2.2436426, 48.8577043),
    (2.2843926, 48.896382), (2.3172704, 48.9088622),
    (2.4006226, 48.9055141), (2.4186822, 48.8829853),
    (2.4182191, 48.8238749), (2.3520005, 48.8058842),
    (2.245958, 48.8311913),
])

client = Client()
dataset = client.dataset(&quot;open_data.copernicus.sentinel2_msi&quot;)
data = dataset.query(
    collections=[&quot;S2A_S2MSI2A&quot;, &quot;S2B_S2MSI2A&quot;, &quot;S2C_S2MSI2A&quot;],
    temporal_extent=(&quot;2025-09-01&quot;, &quot;2025-10-01&quot;),
    spatial_extent=paris,
    show_progress=True,
)
print(data)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;No API key required. The &lt;code&gt;Client()&lt;/code&gt; constructor connects to the Tilebox open data catalog by default when querying public datasets. The &lt;code&gt;dataset.query()&lt;/code&gt; call queries across all three Sentinel-2 satellites (2A, 2B, 2C) and returns the results as an &lt;code&gt;xarray.Dataset&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The spatial extent here is a polygon around central Paris, but it accepts any &lt;code&gt;Polygon&lt;/code&gt; or &lt;code&gt;MultiPolygon&lt;/code&gt; — country outlines, watersheds, coastlines, or any arbitrary shape you can define with Shapely. Tilebox also handles antimeridian crossings and &lt;a href=&quot;https://docs.tilebox.com/datasets/geometries#pole-coverings&quot;&gt;pole-covering geometries&lt;/a&gt; correctly out of the box.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;&amp;lt;xarray.Dataset&amp;gt; Size: 19kB
Dimensions:                (time: 42)
Coordinates:
  * time                   (time) datetime64[ns] ...
Data variables: (12/23)
    id                     (time) &amp;lt;U36 ...
    geometry               (time) object ...
    granule_name           (time) object ...
    cloud_cover            (time) float64 ...
    thumbnail              (time) object ...
    ...&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Each row is one Sentinel-2 scene. The dataset includes the scene geometry, cloud cover percentage, thumbnail URL, granule name, and 20+ other metadata fields — all indexed by acquisition time.&lt;/p&gt;
&lt;h2 id=&quot;explore-the-results-with-xarray&quot;&gt;Explore the results with xarray&lt;/h2&gt;
&lt;p&gt;The output is a standard xarray Dataset. You can filter, slice, and analyze it with any operation xarray supports.&lt;/p&gt;
&lt;p&gt;Check how many scenes were acquired and what the cloud cover distribution looks like:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;print(f&quot;Found {data.sizes[&apos;time&apos;]} Sentinel-2 scenes over Paris in September 2025&quot;)
print(f&quot;Cloud cover: min={data.cloud_cover.min().item():.1f}%, &quot;
      f&quot;max={data.cloud_cover.max().item():.1f}%, &quot;
      f&quot;mean={data.cloud_cover.mean().item():.1f}%&quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code&gt;Found 42 Sentinel-2 scenes over Berlin in September 2025
Cloud cover: min=0.2%, max=99.8%, mean=48.3%&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;filter-by-cloud-cover&quot;&gt;Filter by cloud cover&lt;/h2&gt;
&lt;p&gt;Most workflows start with filtering out cloudy scenes. xarray makes this straightforward:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;clear_scenes = data.where(data.cloud_cover &amp;lt; 20, drop=True)
print(f&quot;{clear_scenes.sizes[&apos;time&apos;]} scenes with less than 20% cloud cover&quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You can also sort by cloud cover to find the clearest acquisitions:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;import numpy as np

sorted_indices = np.argsort(data.cloud_cover.values)
clearest = data.isel(time=sorted_indices[:5])
for i in range(clearest.sizes[&apos;time&apos;]):
    scene = clearest.isel(time=i)
    print(f&quot;{scene.time.dt.strftime(&apos;%Y-%m-%d&apos;).item()} — &quot;
          f&quot;{scene.cloud_cover.item():.1f}% cloud cover — &quot;
          f&quot;{scene.granule_name.item()}&quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;plot-scene-footprints-on-an-interactive-map&quot;&gt;Plot scene footprints on an interactive map&lt;/h2&gt;
&lt;p&gt;Install the visualization dependencies:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;uv add folium matplotlib    # or: pip install folium matplotlib&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now plot every scene footprint on a folium map, colored by cloud cover percentage. Green means clear skies, red means heavy cloud cover.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;import folium
import matplotlib
import matplotlib.cm as cm

# Create a map centered on our area of interest
center_lat = (52.3 + 52.7) / 2
center_lon = (13.1 + 13.8) / 2
m = folium.Map(location=[center_lat, center_lon], zoom_start=10)

# Color scale: green (clear) to red (cloudy)
norm = matplotlib.colors.Normalize(vmin=0, vmax=100)
colormap = cm.RdYlGn_r

for i in range(data.sizes[&apos;time&apos;]):
    scene = data.isel(time=i)
    cloud = scene.cloud_cover.item()
    rgba = colormap(norm(cloud))
    color = matplotlib.colors.to_hex(rgba)

    # Extract polygon coordinates from the geometry
    geom = scene.geometry.item()
    coords = [(lat, lon) for lon, lat in geom.exterior.coords]

    folium.Polygon(
        locations=coords,
        color=color,
        fill=True,
        fill_color=color,
        fill_opacity=0.4,
        popup=f&quot;{scene.time.dt.strftime(&apos;%Y-%m-%d&apos;).item()}&amp;lt;br&amp;gt;&quot;
              f&quot;Cloud cover: {cloud:.1f}%&amp;lt;br&amp;gt;&quot;
              f&quot;{scene.granule_name.item()}&quot;,
    ).add_to(m)

# Add a bounding box for the area of interest
folium.Rectangle(
    bounds=[(52.3, 13.1), (52.7, 13.8)],
    color=&quot;blue&quot;,
    fill=False,
    weight=2,
).add_to(m)

m&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Open the map in a Jupyter notebook and click on any footprint to see the acquisition date, cloud cover, and granule name.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://tilebox.com/images/publication/d9cedaf61f97f511.webp&quot; alt=&quot;Scene footprint&quot;&gt;&lt;figcaption&gt;Scene footprint&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;preview-a-scene-thumbnail&quot;&gt;Preview a scene thumbnail&lt;/h2&gt;
&lt;p&gt;Each Sentinel-2 scene in the Tilebox catalog includes a thumbnail URL. You can display it directly in a notebook:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;from IPython.display import Image, display

# Pick the clearest scene
sorted_indices = np.argsort(data.cloud_cover.values)
best_scene = data.isel(time=sorted_indices[0])

print(f&quot;Clearest scene: {best_scene.time.dt.strftime(&apos;%Y-%m-%d&apos;).item()}, &quot;
      f&quot;{best_scene.cloud_cover.item():.1f}% cloud cover&quot;)
display(Image(url=best_scene.thumbnail.item(), width=400))&lt;/code&gt;&lt;/pre&gt;
&lt;figure&gt;&lt;img src=&quot;https://tilebox.com/images/publication/36d533d2f4c7b9cd.webp&quot; alt=&quot;Scene thumbnail&quot;&gt;&lt;figcaption&gt;Scene thumbnail&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;what-you-just-did&quot;&gt;What you just did&lt;/h2&gt;
&lt;p&gt;In under 30 lines of code, you:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Connected to Tilebox’s open data catalog — no API key, no account&lt;/li&gt;
&lt;li&gt;Queried all Sentinel-2 Level 2A scenes over Paris for a full month&lt;/li&gt;
&lt;li&gt;Filtered scenes by cloud cover using standard xarray operations&lt;/li&gt;
&lt;li&gt;Plotted every scene footprint on an interactive map, color-coded by cloud cover&lt;/li&gt;
&lt;li&gt;Previewed a scene thumbnail&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The query returned structured metadata as an xarray Dataset: scene geometries, acquisition timestamps, cloud cover percentages, thumbnail URLs, and granule identifiers. From here you can feed the results into any processing pipeline, download the actual data from Copernicus Data Space, or build automated monitoring workflows with Tilebox.&lt;/p&gt;
&lt;h2 id=&quot;downloading-the-actual-scene-data&quot;&gt;Downloading the actual scene data&lt;/h2&gt;
&lt;p&gt;Tilebox indexes satellite metadata — the full scene data (bands, reflectance values) lives in the &lt;a href=&quot;https://dataspace.copernicus.eu/&quot;&gt;Copernicus Data Space Ecosystem&lt;/a&gt;. Once you’ve identified the scenes you want, you can download them using the Tilebox storage client:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;from tilebox.storage import CopernicusStorageClient

storage_client = CopernicusStorageClient(
    access_key=&quot;YOUR_CDSE_ACCESS_KEY&quot;,
    secret_access_key=&quot;YOUR_CDSE_SECRET_KEY&quot;,
)

for i in range(clear_scenes.sizes[&apos;time&apos;]):
    scene = clear_scenes.isel(time=i)
    path = storage_client.download(scene)
    print(f&quot;Downloaded to {path}&quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This requires free &lt;a href=&quot;https://dataspace.copernicus.eu/&quot;&gt;Copernicus Data Space credentials&lt;/a&gt; — but querying and filtering metadata through Tilebox is completely open.&lt;/p&gt;
&lt;h2 id=&quot;full-script&quot;&gt;Full script&lt;/h2&gt;
&lt;p&gt;Here’s the complete script in one block. Copy it into a Jupyter notebook or Python file and run it top to bottom.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;from shapely import Polygon
from tilebox.datasets import Client
import numpy as np
import folium
import matplotlib
import matplotlib.cm as cm
from IPython.display import Image, display

# 1. Define your area of interest
berlin = Polygon((
    (13.1, 52.3),
    (13.1, 52.7),
    (13.8, 52.7),
    (13.8, 52.3),
    (13.1, 52.3),
))

# 2. Query Sentinel-2 scenes (no API key needed)
client = Client()
dataset = client.dataset(&quot;open_data.copernicus.sentinel2_msi&quot;)
data = dataset.query(
    collections=[&quot;S2A_S2MSI2A&quot;, &quot;S2B_S2MSI2A&quot;, &quot;S2C_S2MSI2A&quot;],
    temporal_extent=(&quot;2025-09-01&quot;, &quot;2025-10-01&quot;),
    spatial_extent=berlin,
    show_progress=True,
)

# 3. Print summary
print(f&quot;Found {data.sizes[&apos;time&apos;]} Sentinel-2 scenes&quot;)
print(f&quot;Cloud cover: min={data.cloud_cover.min().item():.1f}%, &quot;
      f&quot;max={data.cloud_cover.max().item():.1f}%, &quot;
      f&quot;mean={data.cloud_cover.mean().item():.1f}%&quot;)

# 4. Filter by cloud cover
clear_scenes = data.where(data.cloud_cover &amp;lt; 20, drop=True)
print(f&quot;{clear_scenes.sizes[&apos;time&apos;]} scenes with less than 20% cloud cover&quot;)

# 5. Plot footprints on an interactive map
center_lat = (52.3 + 52.7) / 2
center_lon = (13.1 + 13.8) / 2
m = folium.Map(location=[center_lat, center_lon], zoom_start=10)

norm = matplotlib.colors.Normalize(vmin=0, vmax=100)
colormap = cm.RdYlGn_r

for i in range(data.sizes[&apos;time&apos;]):
    scene = data.isel(time=i)
    cloud = scene.cloud_cover.item()
    rgba = colormap(norm(cloud))
    color = matplotlib.colors.to_hex(rgba)
    geom = scene.geometry.item()
    coords = [(lat, lon) for lon, lat in geom.exterior.coords]

    folium.Polygon(
        locations=coords,
        color=color,
        fill=True,
        fill_color=color,
        fill_opacity=0.4,
        popup=f&quot;{scene.time.dt.strftime(&apos;%Y-%m-%d&apos;).item()}&amp;lt;br&amp;gt;&quot;
              f&quot;Cloud cover: {cloud:.1f}%&amp;lt;br&amp;gt;&quot;
              f&quot;{scene.granule_name.item()}&quot;,
    ).add_to(m)

folium.Rectangle(
    bounds=[(52.3, 13.1), (52.7, 13.8)],
    color=&quot;blue&quot;,
    fill=False,
    weight=2,
).add_to(m)

m

# Display the map in a Jupyter notebook

# 6. Preview the clearest scene thumbnail
sorted_indices = np.argsort(data.cloud_cover.values)
best_scene = data.isel(time=sorted_indices[0])
print(f&quot;\nClearest scene: {best_scene.time.dt.strftime(&apos;%Y-%m-%d&apos;).item()}, &quot;
      f&quot;{best_scene.cloud_cover.item():.1f}% cloud cover&quot;)
display(Image(url=best_scene.thumbnail.item(), width=400))&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;next-steps&quot;&gt;Next steps&lt;/h2&gt;
&lt;p&gt;To run this yourself, install Tilebox with &lt;code&gt;uv add tilebox&lt;/code&gt; (or &lt;code&gt;pip install tilebox&lt;/code&gt;) and open the full notebook from the &lt;a href=&quot;https://github.com/tilebox/examples&quot;&gt;tilebox/examples&lt;/a&gt; repository. Swap out the Paris polygon for your own coordinates and you’ll have a map of your area’s Sentinel-2 coverage in under five minutes.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/3583041da4f66b08.webp" medium="image"><media:title>How to Query Sentinel-2 Data in Python (No Setup Required)</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/3583041da4f66b08.webp"/><category>article</category><pubDate>Mon, 06 Apr 2026 00:00:00 GMT</pubDate><dc:creator>Stefan Amberger</dc:creator></item><item><title>Connecting AI Assistants with Tilebox: Grounding LLMs in Live Space Data</title><link>https://tilebox.com/dispatch/articles/grounding-llms-in-live-space-data</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/grounding-llms-in-live-space-data</guid><description>Learn how the Tilebox MCP Server connects AI assistants to live dataset schemas and workflow statuses, grounding LLM responses in real-time space data.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/220d5179dafb48f4.webp&quot; alt=&quot;Connecting AI Assistants with Tilebox: Grounding LLMs in Live Space Data&quot;&gt;&lt;/p&gt;
&lt;p&gt;Tilebox now supports MCP Servers to bridge the gap between Large Language Models and your dynamic space data. Built on the open Model Context Protocol, this integration allows Large Language Models (LLMs) to interact directly with your live Tilebox environment. This provides responses grounded in your actual data schemas and job statuses.&lt;/p&gt;
&lt;p&gt;Static documentation is often insufficient for high velocity space data engineering. To build reliable pipelines, your AI assistants need more than general knowledge. They require real time context from your specific datasets and workflows.&lt;/p&gt;
&lt;h2 id=&quot;what-is-the-tilebox-mcp-server&quot;&gt;What is the Tilebox MCP Server?&lt;/h2&gt;
&lt;p&gt;The Model Context Protocol (MCP) establishes a standardized connection between AI applications and external data. The Tilebox MCP server exposes tools that enable your assistant to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Introspect Schemas:&lt;/strong&gt; Retrieve the precise, up-to-the-minute schema of your custom datasets (including field types and descriptions).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Query Datasets:&lt;/strong&gt; Assistants can list and filter data points directly from the chat interface.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Monitor Workflows:&lt;/strong&gt; Check the status of Jobs, visualize task dependencies, and identify failures in your Task Runners in real time.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When an AI tool like Claude Code, Cursor, Amp, or OpenAI Codex is configured with the Tilebox MCP server, it can proactively invoke these relevant tools. For example, when you ask a question or request to write code about a specific dataset, the AI tool can query the Tilebox MCP server for the precise, current dataset schema, using that accurate information to generate the code and ground its response.&lt;/p&gt;
&lt;h2 id=&quot;how-the-tilebox-mcp-server-works&quot;&gt;How the Tilebox MCP Server Works&lt;/h2&gt;
&lt;p&gt;Tilebox offers one comprehensive MCP server to cover both operational data (datasets and workflows) and documentation: &amp;lt;&lt;a href=&quot;https://mcp.tilebox.com/mcp&quot;&gt;https://mcp.tilebox.com/mcp&lt;/a&gt;&amp;gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Purpose: Provides tools for accessing and interacting with Tilebox datasets and workflows. It also provides access to the official Tilebox documentation.&lt;/li&gt;
&lt;li&gt;Requirement: Requires authentication using a Tilebox API key.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This single server allows an AI assistant to learn the general syntax and usage of Tilebox APIs and then customize its responses or generated code based on the actual available datasets and their exact schema.&lt;/p&gt;
&lt;h3 id=&quot;customized-coding-assistant&quot;&gt;Customized Coding Assistant&lt;/h3&gt;
&lt;p&gt;A developer is using an AI coding assistant (like Claude or Cursor) to write a script that interacts with a specific Tilebox dataset.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;API Syntax Learning: The assistant uses the docs MCP server (&lt;code&gt;https://docs.tilebox.com/mcp&lt;/code&gt;) to learn the general syntax and usage of the Tilebox APIs.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Crucially, it then queries the data/workflow MCP server (&lt;code&gt;https://mcp.tilebox.com/&lt;/code&gt;) to retrieve the &lt;em&gt;current, precise schema&lt;/em&gt; of the target dataset.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Benefit:&lt;/strong&gt; By grounding documentation syntax with live data schemas, the AI assistant generates code that is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Syntactically correct and follows the latest Tilebox SDK patterns for Go or Python.&lt;/li&gt;
&lt;li&gt;Contextually accurate and uses the exact field names and data types (e.g., cloud_cover, precise_time) from your specific collection.&lt;/li&gt;
&lt;li&gt;Error-proof by preventing common runtime errors caused by typos in field strings or mismatched schema expectations.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;dynamic-and-grounded-qa&quot;&gt;Dynamic and Grounded Q&amp;amp;A&lt;/h3&gt;
&lt;p&gt;An analyst needs a quick answer to a question that requires live data context, such as “What was the total workflow run time yesterday?”&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The AI assistant leverages the data/workflow MCP server’s tools to query the live datasets and workflows in Tilebox.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benefit:&lt;/strong&gt; The AI provides an accurate, up-to-the-minute answer that is explicitly grounded in the live Tilebox data, allowing the user to trust the output without manual verification.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;configuration--setup&quot;&gt;Configuration &amp;amp; Setup&lt;/h2&gt;
&lt;p&gt;You can connect your AI tools to Tilebox using the &lt;code&gt;http&lt;/code&gt; transport.&lt;/p&gt;
&lt;h3 id=&quot;1-json-configuration-snippet&quot;&gt;1. JSON Configuration Snippet&lt;/h3&gt;
&lt;p&gt;For the data/workflow server, remember to include an &lt;code&gt;Authorization&lt;/code&gt; header with your Tilebox API key as the bearer token. Use this for tools like Cursor or the Claude Desktop app. Replace the placeholder with your Tilebox API Key created in the Console.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;{
&amp;nbsp;&amp;nbsp;&quot;mcpServers&quot;: {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;tilebox&quot;: {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;url&quot;: &quot;https://mcp.tilebox.com/&quot;,
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;headers&quot;: {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;Authorization&quot;: &quot;Bearer &amp;lt;YOUR_TILEBOX_API_KEY&amp;gt;&quot;
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;2-for-claude-code-cli&quot;&gt;2. For Claude Code CLI&lt;/h3&gt;
&lt;p&gt;Add these servers directly via your terminal:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;{
&amp;nbsp;&amp;nbsp;&quot;mcpServers&quot;: {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;tilebox&quot;: {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;url&quot;: &quot;https://mcp.tilebox.com/&quot;,
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;headers&quot;: {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&quot;Authorization&quot;: &quot;Bearer &amp;lt;YOUR_TILEBOX_API_KEY&amp;gt;&quot;
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Note: You can still download our full documentation as a markdown file at &lt;a href=&quot;http://docs.tilebox.com/llms-full.txt&quot;&gt;docs.tilebox.com/llms-full.txt&lt;/a&gt; and upload it manually to any LLM.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&quot;https://tilebox.com/images/publication/977927c0f3f69ab2.webp&quot; alt=&quot;&quot;&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=U6VWwtNledQ&quot;&gt;Watch on YouTube&lt;/a&gt;&lt;/figure&gt; 
&lt;p&gt;&lt;em&gt;A practical demo of grounding LLMs in live operational data using MCP&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;mcp-in-action&quot;&gt;MCP in Action&lt;/h2&gt;
&lt;p&gt;The following scenarios illustrate how the Tilebox MCP server transforms your AI assistant into a proactive operations partner.&lt;/p&gt;
&lt;h3 id=&quot;1-precision-spatio-temporal-filtering&quot;&gt;1. Precision Spatio Temporal Filtering&lt;/h3&gt;
&lt;p&gt;Traditional AI assistants often hallucinate field names for geospatial data. With MCP, the assistant queries your live collection schema before writing any code.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;User Prompt:&lt;/strong&gt; “Write a Python script to find all Sentinel 2 granules from last week that fully contain the city of Denver.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Action:&lt;/strong&gt; The assistant uses the Tilebox MCP tool to retrieve the live schema for the Sentinel 2 dataset.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Output:&lt;/strong&gt; It generates a script using &lt;code&gt;spatial_extent&lt;/code&gt; with &lt;code&gt;mode: contains&lt;/code&gt; for a polygon around Denver. This ensures the query is technically valid on the first run.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;2-multi-cluster-job-monitoring&quot;&gt;2. Multi Cluster Job Monitoring&lt;/h3&gt;
&lt;p&gt;Monitoring distributed execution across heterogeneous environments is often complex. The MCP server allows your AI to act as a mission control interface.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;User Prompt:&lt;/strong&gt; “Check the status of the Mosaic job. Are the GPU intensive tasks stuck on the cloud cluster?”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Action:&lt;/strong&gt; The assistant invokes tools to find the Job ID and filters tasks by their assigned &lt;code&gt;cluster_slug&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Output:&lt;/strong&gt; “The root task is running on your on prem cluster. However, 12 subtasks assigned to the gpu cloud cluster are currently in the QUEUED state because no Task Runners are active in that cluster.”&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;3-automated-error-diagnostics&quot;&gt;3. Automated Error Diagnostics&lt;/h3&gt;
&lt;p&gt;When a workflow fails, you can use the AI to identify the exact point of failure without digging through logs manually.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;User Prompt:&lt;/strong&gt; “Why did my last data ingestion job fail?”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Action:&lt;/strong&gt; The assistant queries the most recent Job with a FAILED state and retrieves the error message reported to the orchestrator.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Output:&lt;/strong&gt; “The task LoadCSV failed on Task Runner node 7 with a ValueError: Invalid UUID. One of the IDs in your source file does not match the required schema for the id field.”&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;By adopting an open protocol approach, Tilebox ensures your AI assistant remains a peer in your infrastructure while staying agnostic, scalable, and secure. This is the crucial step in reducing downtime and accelerating payload-to-platform revenue.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/220d5179dafb48f4.webp" medium="image"><media:title>Connecting AI Assistants with Tilebox: Grounding LLMs in Live Space Data</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/220d5179dafb48f4.webp"/><category>article</category><pubDate>Tue, 13 Jan 2026 00:00:00 GMT</pubDate><dc:creator>Stefan Amberger</dc:creator></item><item><title>The 6 Workflow Archetypes: A guide from discovery to services</title><link>https://tilebox.com/dispatch/articles/the-6-workflow-archetypes-a-guide-from-discovery-to-services</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/the-6-workflow-archetypes-a-guide-from-discovery-to-services</guid><description>Six workflow patterns for geospatial data processing with Tilebox, from in-orbit edge computing and near-real-time triggers to batch and scheduled jobs.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/6f114b9be3f72511.webp&quot; alt=&quot;The 6 Workflow Archetypes: A guide from discovery to services&quot;&gt;&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Bottom Line Up Front:&lt;/strong&gt; This post outlines six powerful workflow patterns in Tilebox, from processing data directly in-orbit to cut downlink costs, to building on-demand data products for your customers. Whether you’re developing algorithms, operating in near real-time, or automating routine tasks, these examples demonstrate how to transform raw geospatial data into actionable intelligence more efficiently.&lt;/p&gt;
&lt;p&gt;Your main focus is on delivering products and insights. Tilebox gives you the power to do that, and without frustrating limitations. We’ve compiled a list of our most useful workflow types with examples to help you improve your current processes. Learn how to scale up, expedite, and diversify your products and services to transform raw geospatial data into actionable intelligence.&lt;/p&gt;
&lt;h2 id=&quot;1-in-orbit-workflows-process-data-at-the-edge&quot;&gt;1. In-Orbit Workflows: Process Data at the Edge&lt;/h2&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/09ccca11d26e8e56.webp&quot; alt=&quot;Satellites orbiting around earth, communicating via laser links&quot;&gt;&lt;/p&gt;
&lt;p&gt;Downlinking raw satellite data is like shipping ore from a mine; it’s bulky and expensive. Processing it in-orbit is like refining it at the source and shipping only the gold. This approach dramatically reduces downlink costs, a substantial part of satellite operations.&lt;/p&gt;
&lt;p&gt;For monitoring use cases requiring extremely low latency, or for hyperspectral providers who need to downlink only specific bands, on-orbit processing is a necessity. It moves beyond simple object detection demos to enable operational, decision-making AI in space.&lt;/p&gt;
&lt;h3 id=&quot;tilebox-highlights&quot;&gt;Tilebox Highlights&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Lightweight and Efficient:&lt;/strong&gt; Tilebox is designed to be resource-efficient, consuming minimal power and compute from precious on-orbit resources. Work can be queued and executed only when compute capacity is available, allowing modules to be scaled up or down to fulfill power budgets.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ground-to-Space API Parity:&lt;/strong&gt; The same APIs for workflow orchestration and management are used on-orbit and on the ground. This allows you to create a “digital twin” of your on-orbit processing state for rapid testing, validation, and iteration of algorithms on the ground before uplinking.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Integrated Observability for Attribution:&lt;/strong&gt; Collect critical attribution data and logs from in-orbit processes. This provides a tight feedback loop with ground operations, which is essential for operationalizing AI and decision-making in space.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Parallel Processing in Space:&lt;/strong&gt; Fully utilize on-orbit compute capabilities by running tasks in parallel, ensuring that you get the most out of your hardware.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;what-youll-need&quot;&gt;What you’ll need&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;On-orbit compute hardware with Tilebox deployed (c.f. from one of our pre-integrated compute module partners).&lt;/li&gt;
&lt;li&gt;A ground-based environment for algorithm development and testing.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;next-steps&quot;&gt;Next Steps&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Start developing your algorithms on the hosted version of Tilebox today; they will be fully compatible with the on-board daemon.&lt;/li&gt;
&lt;li&gt;Contact our team to discuss how to implement Tilebox for your on-orbit processing needs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;2-distributed-workflows-process-data-across-multiple-environments&quot;&gt;2. Distributed Workflows: Process Data Across Multiple Environments&lt;/h2&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/b7357691d28972c0.webp&quot; alt=&quot;Workflow visualization with tasks in two clouds and on-premise&quot;&gt;&lt;/p&gt;
&lt;p&gt;As remote sensing data is often distributed across various locations, workflows that span multiple clouds and on-premise compute environments are becoming increasingly common. Tilebox simplifies the implementation of these complex workflows, such as data fusion use-cases like wildfire detection using both Landsat and Sentinel data. All workflows may be distributed.&lt;/p&gt;
&lt;h3 id=&quot;tilebox-highlights-1&quot;&gt;Tilebox Highlights&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Process Data at the Source:&lt;/strong&gt; Minimize costly data egress fees and reduce latency by running compute tasks directly where your data resides, whether on-premise or across multiple cloud providers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Natively Heterogeneous:&lt;/strong&gt; Orchestrate workflows that span on-premise servers, multiple public clouds, and even on-orbit compute, without requiring complex Kubernetes setups.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Utilize Specialized Hardware:&lt;/strong&gt; Route specific tasks to clusters with specialized hardware, such as GPUs or high-memory nodes, ensuring optimal performance for every step of your pipeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Simplified Multi-Cluster Dependencies:&lt;/strong&gt; Declare dependencies between tasks in different environments as easily as if they were on the same machine. Tilebox handles the complex orchestration.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resilience Across Environments:&lt;/strong&gt; The failure of a node in one environment doesn’t cascade. The workflow’s built-in fault tolerance ensures that the overall process can continue even with localized infrastructure issues.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;what-youll-need-1&quot;&gt;What you’ll need&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Access to two or more distinct compute environments (e.g., on-premise and public cloud).&lt;/li&gt;
&lt;li&gt;Tilebox Task Runners deployed in each of these environments.&lt;/li&gt;
&lt;li&gt;Tasks designed to run in specific environments, orchestrated via Tilebox Clusters.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;next-steps-1&quot;&gt;Next Steps&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.tilebox.com/workflows/concepts/clusters&quot;&gt;Read about how to manage clusters&lt;/a&gt; for distributed workflows.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://console.tilebox.com/workflows/clusters&quot;&gt;Create a cluster&lt;/a&gt; in the console&lt;/li&gt;
&lt;li&gt;Start two runners, one in each cluster&lt;/li&gt;
&lt;li&gt;Schedule work from a task running on one cluster to the other&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;3-near-real-time-workflows-trigger-processing-on-incoming-data&quot;&gt;3. Near Real-Time Workflows: Trigger Processing on Incoming Data&lt;/h2&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/5e024e27a3ba568f.webp&quot; alt=&quot;schematic of workflows triggered by satellite downlinks&quot;&gt;&lt;/p&gt;
&lt;p&gt;When speed is critical, workflows can be triggered in near real-time as new data becomes available. This is necessary for applications like wildfire detection or maritime surveillance, which require immediate processing of newly published data.&lt;/p&gt;
&lt;p&gt;For instance, when a satellite downlinks new payload data, it can automatically trigger the L0 to L1 processing chain.&lt;/p&gt;
&lt;h3 id=&quot;tilebox-highlights-2&quot;&gt;Tilebox Highlights&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Low-Overhead Processing:&lt;/strong&gt; The time it takes to start a task is in the millisecond range, making Tilebox highly effective for latency-sensitive applications with frequent, small workloads.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Diverse Triggering Mechanisms:&lt;/strong&gt; Initiate workflows automatically from a variety of sources, including Storage Events in cloud buckets (GCS, S3) and local file systems (both in closed beta)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Robust Backpressure Management:&lt;/strong&gt; Gracefully handles sudden bursts of trigger events, ensuring system stability and preventing your infrastructure from being overloaded during data storms.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Full Observability:&lt;/strong&gt; Monitor the health and performance of your automated pipelines with integrated Tracing and Logging, which is essential for production environments.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;what-youll-need-2&quot;&gt;What you’ll need&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;An event source configured to trigger workflows (e.g., a cloud storage bucket with notifications enabled).&lt;/li&gt;
&lt;li&gt;A Task defined to handle the incoming event data.&lt;/li&gt;
&lt;li&gt;A continuously running Tilebox Task Runner listening for these jobs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;next-steps-2&quot;&gt;Next Steps&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.tilebox.com/workflows/automations/storage-events&quot;&gt;Learn how to set up automations&lt;/a&gt; to trigger workflows from events.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;4-customer-triggered-workflows-enable-on-demand-processing&quot;&gt;4. Customer-Triggered Workflows: Enable On-Demand Processing&lt;/h2&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/41d49894ebc7a2c4.webp&quot; alt=&quot;Workflow tasks triggered and processed as new data arrives over time&quot;&gt;&lt;/p&gt;
&lt;p&gt;Empower your customers by allowing them to trigger asynchronous jobs directly from a customer portal or API. This is ideal for generating custom reports, visualizations, or performing exploratory processing of specific areas of interest.&lt;/p&gt;
&lt;h3 id=&quot;tilebox-highlights-3&quot;&gt;Tilebox Highlights&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;API-Driven Automation:&lt;/strong&gt; Expose your processing capabilities directly to customers through a secure API, allowing them to request custom products on-demand.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Asynchronous Execution:&lt;/strong&gt; Jobs run in the background, allowing your customer-facing portal or API to remain responsive. You can notify customers upon completion.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Secure by Design:&lt;/strong&gt; The underlying architecture allows compute nodes to run in isolated environments without internet exposure, protecting your core infrastructure and algorithms.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;what-youll-need-3&quot;&gt;What you’ll need&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;A customer-facing application (e.g., web portal, mobile app) with an API.&lt;/li&gt;
&lt;li&gt;Backend logic that translates customer requests into Tilebox API calls to submit jobs.&lt;/li&gt;
&lt;li&gt;A scalable set of Task Runners to handle potentially bursty customer demand.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;next-steps-3&quot;&gt;Next Steps&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Expose an API endpoint that your customers can use to trigger a Tilebox workflow.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;5-batch-workflows-develop-and-backtest-algorithms-at-scale&quot;&gt;5. Batch Workflows: Develop and Backtest Algorithms at Scale&lt;/h2&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/2ad0adcac218e4a5.webp&quot; alt=&quot;A large task tree illustrating jobs with millions of tasks&quot;&gt;&lt;/p&gt;
&lt;p&gt;Batch processing is fundamental for developing and backtesting algorithms. It allows you to reprocess entire missions or apply new algorithms to large volumes of historical data, which is critical for creating and validating new products.&lt;/p&gt;
&lt;p&gt;These workflows are ideal for large-scale processing that covers vast areas of interest or long time periods, such as creating quarterly cloud-free mosaics or running crop yield estimates.&lt;/p&gt;
&lt;h3 id=&quot;tilebox-highlights-4&quot;&gt;Tilebox Highlights&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dynamic Workflows:&lt;/strong&gt; Unlike systems that use static DAGs, a Tilebox workflow can adapt at runtime. Tasks can determine subsequent steps based on intermediate results, enabling more complex and data-dependent processing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost-Efficient Retries:&lt;/strong&gt; For long-running processes, failures can be expensive. Tilebox allows workflows to be restarted from the point of failure after a bug is fixed, saving significant time and infrastructure costs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Simplified Infrastructure:&lt;/strong&gt; Tilebox simplifies deployment by not requiring complex cluster managers like Kubernetes. Since compute nodes operate independently without direct communication, network setup is significantly easier.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flexibility and Scalability:&lt;/strong&gt; Implement tasks in your preferred language—not just Python—and scale your jobs to handle millions of tasks.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;what-youll-need-4&quot;&gt;What you’ll need&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;An algorithm or processing logic defined in Tilebox Tasks.&lt;/li&gt;
&lt;li&gt;A dataset to process (e.g., a large data archive or a custom Tilebox Dataset).&lt;/li&gt;
&lt;li&gt;Compute infrastructure (even a local machine) to run a Tilebox Task Runner.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;next-steps-4&quot;&gt;Next Steps&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/tilebox/examples/tree/main/workflows-hello-world-py&quot;&gt;Explore our workflows hello world&lt;/a&gt; to get started with your first batch workflow.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.google.com/search?q=https://docs.tilebox.com/workflows/concepts/jobs%23submission&quot;&gt;Read the documentation on submitting jobs&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;6-scheduled-workflows-automate-repetitive-tasks&quot;&gt;6. Scheduled Workflows: Automate Repetitive Tasks&lt;/h2&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/e934d3e32f14d41d.webp&quot; alt=&quot;Workflow task trees created successively over time for on-demand processing&quot;&gt;&lt;/p&gt;
&lt;p&gt;Automate recurring tasks to ensure consistency and reduce manual effort. This is perfect for operations that need to run at regular intervals, such as scraping tasks that operate hourly or workflows that generate daily mosaics and weekly reports.&lt;/p&gt;
&lt;h3 id=&quot;tilebox-highlights-5&quot;&gt;Tilebox Highlights&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;First-Class CRON Support:&lt;/strong&gt; Scheduled automations are a core feature, allowing you to reliably run tasks at any interval.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Seamless Updates with Versioning:&lt;/strong&gt; Deploy updated task logic at any time. The next scheduled run will automatically use the new version, enabling rolling updates to your jobs without downtime or manual changes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flexible Management:&lt;/strong&gt; Manage your scheduled automations programmatically via the SDK or through the intuitive Tilebox Console.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Code Reusability:&lt;/strong&gt; The framework for batch and scheduled workflows is the same, allowing for quick operationalization by reusing code between different modes.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;what-youll-need-5&quot;&gt;What you’ll need&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;A recurring task defined as a Tilebox Task (e.g., generating a daily report).&lt;/li&gt;
&lt;li&gt;An automation configured with a CRON schedule in the Tilebox Console or via the SDK.&lt;/li&gt;
&lt;li&gt;A continuously running Tilebox Task Runner to execute the scheduled jobs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;next-steps-5&quot;&gt;Next Steps&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.tilebox.com/workflows/automations/cron&quot;&gt;See how to configure CRON triggers&lt;/a&gt; in the documentation.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/tilebox/examples/tree/main/workflows-cron-automation-py&quot;&gt;Explore our CRON example&lt;/a&gt; on Github&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Want to explore something else? &lt;a href=&quot;https://discord.gg/Rd4MFYdnht&quot;&gt;Join us on Discord&lt;/a&gt; for community support, ideation, and more.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/6f114b9be3f72511.webp" medium="image"><media:title>The 6 Workflow Archetypes: A guide from discovery to services</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/6f114b9be3f72511.webp"/><category>article</category><pubDate>Mon, 25 Aug 2025 00:00:00 GMT</pubDate><dc:creator>Stefan Amberger</dc:creator></item><item><title>Cloud-Free, Country-Scale Mosaic in Under 3 Hours: A Tilebox Workflow</title><link>https://tilebox.com/dispatch/articles/cloud-free-mosaic</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/cloud-free-mosaic</guid><description>How we built a cloud-free Sentinel-2 mosaic over Ireland using Tilebox Workflows with multi-environment execution, parallel Zarr writes, on 700 granules.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/4e45eb0b1316b53b.webp&quot; alt=&quot;Cloud-Free, Country-Scale Mosaic in Under 3 Hours: A Tilebox Workflow&quot;&gt;&lt;/p&gt;
&lt;p&gt;Creating large-scale, cloud-free satellite imagery mosaics demands efficient data handling and powerful processing. Inspired by a data scientist’s approach to generating a quarterly cloud-free Sentinel-2 10m resolution RGB mosaic over Ireland, we replicated this process using Tilebox to demonstrate how use cases like this can be streamlined. This post highlights how Tilebox streamlines complex geospatial workflows, leveraging multi-environment execution and parallel writes to a Zarr datacube for optimal performance.&lt;/p&gt;
&lt;p&gt;Our workflow involves four key steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Locating relevant Sentinel-2 input granules over Ireland within a three-month period&lt;/li&gt;
&lt;li&gt;Reading the Red, Green, and Blue (RGB) bands, plus the scene classification layer (cloud mask) from each located Sentinel-2 product&lt;/li&gt;
&lt;li&gt;Reprojecting every product onto a common grid&lt;/li&gt;
&lt;li&gt;Aggregating data across the time dimension to produce a single cloud-free measurement for every pixel&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If you want to skip ahead and see our results for yourself, check out the &lt;a href=&quot;https://tilebox.com/dispatch/articles/cloud-free-mosaic#visualizing-the-results&quot;&gt;interactive visualization down below.&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;efficient-granule-discovery-with-tilebox-datasets&quot;&gt;Efficient Granule Discovery with Tilebox Datasets&lt;/h2&gt;
&lt;p&gt;The first step, locating the necessary Sentinel-2 granules, is straightforward with Tilebox’s spatio-temporal query capabilities. Our tilebox.datasets client allows for rapid discovery of relevant data.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;from tilebox.datasets import Client
from shapely.geometry import box

# Initialize the Tilebox datasets client
client = Client()
sentinel_2 = client.dataset(&quot;open_data.copernicus.sentinel2_msi&quot;)
sentinel_2a = sentinel_2.collection(&quot;S2A_S2MSI2A&quot;)

# Define a rectangular area over Ireland
area = box(-10.68234795, 51.36473433, -5.34679566, 55.44704815)

# Query for Sentinel-2 granules within the specified temporal and spatial extent
granules = sentinel_2a.query(
    temporal_extent=(&quot;2025-03-01&quot;, &quot;2025-06-01&quot;),
    spatial_extent=area
)

print(f&quot;Located {granules.sizes[&apos;time&apos;]} Sentinel-2 granules.&quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In just milliseconds, we located around 700 Sentinel-2A granules needed for our mosaic. &lt;a href=&quot;https://docs.tilebox.com/datasets/open-data&quot;&gt;Tilebox Open Data&lt;/a&gt; provides instant access to public datasets, and its robust spatial indexing handles complex geometries like antimeridian crossings seamlessly.&lt;/p&gt;
&lt;h2 id=&quot;managing-data-volumes&quot;&gt;Managing Data Volumes&lt;/h2&gt;
&lt;p&gt;For our true-color mosaic, we need the Red, Green, and Blue bands at 10m resolution, plus the 20m resolution cloud mask for filtering. A single Sentinel-2 L2A granule, with these four bands, amounts to around 315MB. The 700 granules we located amount to a total of approximately 155GB of data. Efficiently handling this volume requires a smart approach to storage and processing.&lt;/p&gt;
&lt;p&gt;To avoid costly data transfers and manage large intermediate products, we chose Zarr as our intermediate storage format. Zarr is a highly efficient, chunked array format ideal for parallel I/O and cloud-native workflows. It allows us to persist reprojected data in a way that supports easy spatial chunking and parallel access across the time dimension.&lt;/p&gt;
&lt;p&gt;We initialized an empty Zarr cube for each band, with dimensions corresponding to the full spatial extent over our area over Ireland and a time dimension corresponding to one time layer for each of our 700 products. This resulted in a data cube shape of time=716, y=37151, x=45419. By setting the time dimension chunk size to 1, we enable parallel writing of individual timestamps. Furthermore, configuring a spatial chunk size (in our case we chose 2048x2048 pixels) allows Zarr to automatically skip writing empty chunks for each time layer, providing immediate efficiency gains, especially when reprojecting smaller granules onto a large target grid. Additionally this is also what allows us to process individual, smaller spatial chunks across the entire time dimension in the required subsequent temporal aggregation.&lt;/p&gt;
&lt;h2 id=&quot;orchestrating-multi-environment-workflows&quot;&gt;Orchestrating Multi-Environment Workflows&lt;/h2&gt;
&lt;p&gt;The process of reading, reprojecting, and writing each Sentinel-2 product to a Zarr cube is inherently parallel. Tilebox Workflows are designed to leverage this parallelism by allowing us to define individual tasks that can be automatically parallelized. All that is required to achieve that is to transform our processing logic into &lt;a href=&quot;https://docs.tilebox.com/workflows/concepts/tasks&quot;&gt;tasks&lt;/a&gt;.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;import rasterio
from tilebox.workflows import Task, ExecutionContext

class GranuleProductToZarr(Task):
    &quot;&quot;&quot;
    Processes a single Sentinel-2 product (e.g., B02, B03, B04, SCL)
    reprojects it, and writes it to the Zarr datacube at a specific time index.
    &quot;&quot;&quot;

    product_location: str
    &quot;&quot;&quot;A concrete Sentinel 2 product to convert to Zarr&quot;&quot;&quot;

    time_index: int
    &quot;&quot;&quot;The time index of the granule in the output Zarr datacube&quot;&quot;&quot;

    def execute(self, context: ExecutionContext) -&amp;gt; None:
        # Open the JPEG2000 product using rasterio
        with rasterio.open(self.product_location, driver=&quot;JP2OpenJPEG&quot;) as product:
             with rasterio.open(self.product_location, driver=&quot;JP2OpenJPEG&quot;) as product:
             arr = product.read(1)
             src_grid = GeoBox(shape=arr.shape, affine=product.transform, crs=product.crs)

         # .... continue with reprojecting the array, and writing it to zarr at the given time index

class GranuleToZarr(Task):
    &quot;&quot;&quot;
    Orchestrates the processing of a single Sentinel-2 granule by
    submitting subtasks for each relevant band.
    &quot;&quot;&quot;
    granule_location: str
    time_index: int

    def execute(self, context: ExecutionContext) -&amp;gt; None:
        products = list_products(  # glob the given directory for objects matching a pattern
            self.granule_location,
            filter=[&quot;R10m/*B02*&quot;, &quot;R10m/*B03*&quot;, &quot;R10m/*B04*&quot;, &quot;R20m/*SCL*&quot;]
        )
        for product in products:
            context.submit_subtask(product, self.time_index)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;code&gt;GranuleToZarr&lt;/code&gt; task spawns multiple &lt;code&gt;GranuleProductToZarr&lt;/code&gt; subtasks, one for each band. These subtasks can then execute in parallel on available task runners. A higher-level task then iterates through the list of all 700 Sentinel-2 granules in a similar fashion, submitting a &lt;code&gt;GranuleToZarr&lt;/code&gt; task for each.&lt;/p&gt;
&lt;h2 id=&quot;co-locating-compute-with-multi-environment-capabilities&quot;&gt;Co-locating Compute with Multi-Environment Capabilities&lt;/h2&gt;
&lt;p&gt;As you may have noticed, the above &lt;code&gt;GranuleToZarr&lt;/code&gt; task assumes Sentinel-2 product files are available as part of the local file system. Traditionally, we would need to adapt this logic to also support reading products via an S3 compatible object store interface. However, a significant advantage of Tilebox is its ability to run workflows across diverse compute environments. To avoid downloading 155GB of Sentinel-2 data from the Copernicus archive (hosted on CloudFerro), we instead executed our workflow in a multi-environment fashion.&lt;/p&gt;
&lt;p&gt;Distributing work across various compute environments is a built-in feature of our workflow orchestrator, so the only requirement for setting this up was to start task runners in the right locations, no adaptations to the source code are necessary.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Access &amp;amp; Reprojection (CloudFerro):&lt;/strong&gt; We deployed a Tilebox task runner on a virtual machine within CloudFerro’s infrastructure. This runner’s sole responsibility is to read Copernicus data, reproject it, and write the intermediate Zarr cube directly to Google Cloud Storage. This minimizes egress from CloudFerro. And it allows us to read the Sentinel products directly from a filesystem, since VMs on CloudFerro have the whole Copernicus archive mounted at &lt;code&gt;/eodata&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Temporal Aggregation (Local Cluster):&lt;/strong&gt; For the final temporal aggregation (producing the cloud-free mosaic), we utilized a makeshift cluster using three of our developer MacBooks. Each MacBook’s task runner fetches a specific 2048x2048 spatial chunk of the Zarr cube across the entire time range, performs the aggregation, and writes the output for that chunk back into the final mosaic layer of our Zarr store. This is made possible by Zarr’s powerful rechunking capabilities and Tilebox’s flexible task distribution. This strategy enabled us to avoid setting up multiple expensive VMs on Cloudferro for compute. Alternatively, we could have just as well used cheap spot instances, which would be the ideal solution for even larger processings, such as a mosaic of the entire globe, and to minimize AWS egress cost.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This multi-environment approach demonstrates Tilebox’s flexibility: seamlessly integrating specialized compute resources (CloudFerro for data proximity, MacBooks for distributed final processing) into a single, cohesive workflow without rewriting code.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/4dd65b853d97cf4f.svg&quot; alt=&quot;Architecture of the distributed Sentinel-2 mosaic workflow&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Figure 1: Architecture diagram depicting a multi-environment workflow, and Zarr chunking mechanics to enable parallelization. Depicted are task runners deployed on a CloudFerro VM as well as locally on developer notebooks.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;visualizing-the-results&quot;&gt;Visualizing the Results&lt;/h2&gt;
&lt;p&gt;The final output is a mosaic, which we converted from Zarr to GeoTIFF and then uploaded to Ellipsis Drive for visualization.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://tilebox.com/dispatch/articles/cloud-free-mosaic#visualizing-the-results&quot;&gt;Explore the interactive mosaic on Tilebox&lt;/a&gt;&lt;/p&gt; 
&lt;p&gt;This workflow showcases how Tilebox simplifies complex, large-scale geospatial processing by combining efficient data access, parallel processing with Zarr, and flexible multi-environment workflow orchestration. The full code for this example is available on &lt;a href=&quot;https://github.com/tilebox/examples/tree/main/s2-cloudfree-mosaic&quot;&gt;GitHub&lt;/a&gt;, and you can try it out yourself using our &lt;a href=&quot;https://tilebox.com/pricing&quot;&gt;free Community Access&lt;/a&gt; tier.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/4e45eb0b1316b53b.webp" medium="image"><media:title>Cloud-Free, Country-Scale Mosaic in Under 3 Hours: A Tilebox Workflow</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/4e45eb0b1316b53b.webp"/><category>article</category><pubDate>Tue, 15 Jul 2025 00:00:00 GMT</pubDate><dc:creator>Lukas Bindreiter</dc:creator></item><item><title>Skip the Respawn Delay: Space Data Workflows with Built-in Resilience</title><link>https://tilebox.com/dispatch/articles/skip-the-respawn</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/skip-the-respawn</guid><description>How Tilebox Workflows handle processing failures with built-in resilience through natural checkpoints, re-entrant execution, and versioned task rollouts.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/d2035463cd78a8c0.webp&quot; alt=&quot;Skip the Respawn Delay: Space Data Workflows with Built-in Resilience&quot;&gt;&lt;/p&gt;&lt;h2 id=&quot;natural-checkpoints-versioned-tasking-and-other-smart-workflow-design-decisions&quot;&gt;Natural checkpoints, versioned tasking, and other smart workflow design decisions&lt;/h2&gt;
&lt;p&gt;During large processing workflows, things can and will go wrong. Memory errors, missing data, network issues, just to name a few. These processing errors cost time and money, resulting in inefficient compute resource utilization, reduced developer productivity, and increased egress costs from redundantly fetched external data dependencies.&lt;/p&gt;
&lt;p&gt;Many workflow orchestrators don’t account for re-entrant execution at all. If something goes wrong, the whole processing has to be re-done. Or, a new workflow has to be developed for every failure that is able to resume from that particular partial state, and explicit checkpoints have to be added to the workflow to persist partial states. Not only does this take valuable time, it should be retired as a historical problem – there is no need to settle for this model.&lt;/p&gt;
&lt;p&gt;Accounting for all possible errors beforehand is not possible, so it’s vital to be ready to act with the right resources to thoroughly investigate, rewrite, and deploy.&lt;/p&gt;
&lt;p&gt;This is precisely why we designed our workflow orchestrator differently. &lt;a href=&quot;https://docs.tilebox.com/workflows/introduction&quot;&gt;Tilebox Tasks&lt;/a&gt; form natural checkpoints, and the idempotent nature of them enables efficient re-entrant execution. If something goes wrong, full observability enables developers to quickly dive in and figure out the problem. And it’s all ready-to-go out of the box.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/6e87d98f6f1beb7b.svg&quot; alt=&quot;Workflow timeline showing two compute nodes and a failure during processing&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;A long running workflow may fail after a lot of the processing work is already completed.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/583c1597e2094096.svg&quot; alt=&quot;Workflow timeline restarting all tasks after a failure&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;One approach to recover from such failures is to just re-run the whole processing, essentially duplicating all the work that has already been computed.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/26ee0a4e248f970f.svg&quot; alt=&quot;A separate recovery workflow finds processed data and resumes the unfinished parts&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Another approach to handle such failures is to develop a new workflow capable of picking up partial results from a previously failed run. While this preserves compute resources, it does take valuable development time.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/fb661a6108d48eeb.svg&quot; alt=&quot;Re-entrant execution resumes unfinished processing tasks and merges the results&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Tilebox re-entrant execution eliminates the necessity for manual recovery tasks and reduces downtime while simplifying the workflow.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;After the issue is identified and fixed, versioned tasks allow gradual rollout without disrupting existing clusters / workflows. Then the workflow orchestrator is able to resume the failed job from the state it was in previously to the failure. Re-entrant processing can be manually triggered or configured using automated retries, giving developers another layer of operational resiliency.&lt;/p&gt;
&lt;p&gt;Join us on &lt;a href=&quot;https://discord.com/invite/Rd4MFYdnht&quot;&gt;discord&lt;/a&gt; for the latest releases, challenges, and Tilebox tips.&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/d2035463cd78a8c0.webp" medium="image"><media:title>Skip the Respawn Delay: Space Data Workflows with Built-in Resilience</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/d2035463cd78a8c0.webp"/><category>article</category><pubDate>Tue, 25 Mar 2025 00:00:00 GMT</pubDate><dc:creator>Stefan Amberger</dc:creator></item><item><title>Beyond General-Purpose: The Need for Space-Data Native Frameworks</title><link>https://tilebox.com/dispatch/articles/space-data-native-framework</link><guid isPermaLink="true">https://tilebox.com/dispatch/articles/space-data-native-framework</guid><description>Why general-purpose tools fall short for satellite data pipelines and how a space-data native framework addresses resilience, scalability, and performance.</description><content:encoded>&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/3868f64f570fb60f.webp&quot; alt=&quot;Beyond General-Purpose: The Need for Space-Data Native Frameworks&quot;&gt;&lt;/p&gt;&lt;p&gt;In-house space data pipelines are finicky, to say the least. A patchwork that is often rebuilt for every major upgrade. Workflows that are impossible to iterate because they are baked into Infrastructure as Code. Infrastructure redundancies that consume unnecessary time and resources, like fixed cluster sizes, the overhead of booting up Docker containers for every task, or difficulties in rolling out updates while keeping production up. The inefficiencies are many, but they are all a result of incompatible design across environments, temporary fixes, and data type limitations.&lt;/p&gt;
&lt;p&gt;There has been no standard or best practices for space data pipelines. To reach profitability and meet the expectations of real world applications it needs a &lt;em&gt;space-data native framework.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Here’s how we designed ours:&lt;/p&gt;
&lt;h2 id=&quot;datasets&quot;&gt;Datasets&lt;/h2&gt;
&lt;p&gt;Let’s start with the critical foundation of data access and storage. Common database models such as Postgres and even specialized time-series databases like InfluxDB are amazing pieces of engineering, but on their own come with constraints that become real obstacles for space data pipelines, particularly regarding Earth Observation. One lacks support for textual data, the other supports geometries but has limited support for spatial indexing, a critical capability of any data catalog, and a very common reason for performance limitations of the data catalogs of the world.&lt;/p&gt;
&lt;p&gt;Instead of wasting time waiting for results or consuming all your CPU on table scans, a framework designed intentionally for these pipelines needs to have:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Support for real-time metadata indexing with customizable data types&lt;/li&gt;
&lt;li&gt;Support for arbitrary data types, including strings, polygons, coordinates&lt;/li&gt;
&lt;li&gt;Fast, reliable spatio-temporal indexing&lt;/li&gt;
&lt;li&gt;A very high performance API to support large queries of telemetry, metadata, or lightweight payload data&lt;/li&gt;
&lt;li&gt;Backwards compatible typing and customizable datasets, or catalogs&lt;/li&gt;
&lt;li&gt;Standards compliant (STAC) output interfaces where desired&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Tilebox affords engineers the freedom to create and edit their own data types – instead of the database schema – without breaking existing software and datasets. Spatio-temporal queries are orders of magnitudes faster than comparable systems. This is essential for streamlined scalability. As operations expand, so will capabilities.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://tilebox.com/images/publication/00e2f0ddb4b8012f.webp&quot; alt=&quot;Spatial-temporal query performance comparison showing 130× and 11× speedups&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;workflows&quot;&gt;Workflows&lt;/h2&gt;
&lt;p&gt;While many companies are still figuring it out on the ground, others are preparing for edge computing. Any space-data native framework has to work in orbit. And, we think, everywhere else. Across clouds, mixed environments, and multiple clusters. Data fusion requires this level of flexibility, plus it empowers any team to work with the exact data they need no matter where it’s stored. This expedites processing and analysis, which creates increased revenue opportunities for time-sensitive data.&lt;/p&gt;
&lt;p&gt;Efficiency is a key requirement for space data, as its source is far away and its volume is massive. Reducing downlink, transfer, and storage costs is one side of the equation, the other is job resilience. Executing workflows on Spot Instances is a necessary cost-savings for large-scale satellite data processing. But what happens when that Spot Instance goes down? What if a job breaks?&lt;/p&gt;
&lt;p&gt;Resiliency affects everything: compute costs, latency, data security.&lt;/p&gt;
&lt;p&gt;Manually developed monitoring tools are one of those inefficient redundancies software teams shouldn’t be wasting their time on building and managing. Even Tilebox leverages a service for this, Axiom.co; pre-integrated with Tilebox to deliver the fastest, most thorough distributed observability across your workflows, with a generic OpenTelemetry exporter available as well.&lt;/p&gt;
&lt;p&gt;Identifying the break is the first step, re-entrant processing is the second step. Tilebox saves all work up to the breakpoint and automatically initiates a retry, keeping you up and running without manual intervention and rescheduling. And for large workflows, Tilebox supports auto-scaling clusters.&lt;/p&gt;
&lt;p&gt;These are not conveniences, they are vital functionalities for a successful space data pipeline.&lt;/p&gt;
&lt;h2 id=&quot;future-proof-your-pipeline&quot;&gt;Future-Proof Your Pipeline&lt;/h2&gt;
&lt;p&gt;The right tooling makes all the difference. Space data engineers need flexibility, resilience, and purpose-built functionality to stay focused on delivering rather than pipeline maintenance and repair. We are building Tilebox for you and the future of space data services. So, create an account and explore our work – and give us feedback on your ideal tooling. It’s time to evolve space data management with high-performance software made for space data.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Join our&lt;/em&gt; &lt;a href=&quot;https://discord.com/invite/Rd4MFYdnht&quot;&gt;&lt;em&gt;discord community&lt;/em&gt;&lt;/a&gt; &lt;em&gt;for updates on our latest releases.&lt;/em&gt;&lt;/p&gt;</content:encoded><media:content url="https://tilebox.com/images/publication/3868f64f570fb60f.webp" medium="image"><media:title>Beyond General-Purpose: The Need for Space-Data Native Frameworks</media:title></media:content><media:thumbnail url="https://tilebox.com/images/publication/3868f64f570fb60f.webp"/><category>article</category><pubDate>Mon, 17 Feb 2025 00:00:00 GMT</pubDate><dc:creator>Megan Van Patten</dc:creator></item></channel></rss>