wikidata: orchestrate spark jobs for data preparation and transfer.

Why:

In order to generate split graphs from the ttl dump, we need to orchestrate spark jobs for data preparation and transfer. Ultimately we need to generate snapshots in n3 format that will be ingested by the T428235: [NEEDS GROOMING] data-reload: create a QLever index using an ephemeral instance.

What:

Ported ETL logic from the Search import_ttl pipeline. Adds support for reading dumps from S3, and transfering n3 files to S3.

TODO:

  • Bikeshedding output paths
  • Add test fixtures

cc @lerickson @trueg

Bug: T428237

Edited by Gmodena

Merge request reports

Loading