BigQuery Loads, Now on Arrow Time

BigQuery Loads, Now on Arrow Time
Join our mailing list

Stay current with all things Meltano

What Changed Under the Hood

Previously, target-bigquery worked row by row, translating each Singer record into BigQuery-ready data before it could load. That works, but it’s not the fastest way to move large volumes of data.

Now, target-bigquery can accept BATCH messages pointing to Arrow IPC files, a columnar format built for exactly this kind of high-throughput handoff. When those messages show up, the target automatically routes them through a new, faster code path. No flags, no settings, no extra steps.

The Numbers

Benchmarks comparing row-by-row RECORD ingestion against the new ARROW+BATCH path:


That’s the difference between a pipeline that’s sweating its SLA and one with room to spare.


Why This Matters for Your Pipelines

  • Speed where it counts. Columnar batch loading avoids the overhead of row-by-row processing, translating directly into the throughput gains above.
  • Nothing new to configure. If your extractor emits Arrow-encoded BATCH messages, target-bigquery just picks them up. Your existing schema-driven table setup stays exactly the same: Arrow data flows into the columns your Singer SCHEMA message already defines.
  • One less thing to manage. Manifest files are treated as consume-once, so you don’t need to worry about cleanup or leftover state after a load.


Built on a Growing Pattern

This isn’t a one-off. Arrow BATCH support is becoming a shared foundation across Meltano targets, so if you’re running a broader pipeline, faster ingestion isn’t limited to just BigQuery.


Try It Out

If you’re already running target-bigquery, this speedup comes to you automatically the next time your extractor sends Arrow-encoded batches, no migration required. Haven’t tried Meltano yet? It might be a good time to see what a modern, config-light data pipeline feels like.

Intrigued?

You haven’t seen nothing yet!