Jobs in Germany

Home  | English Speaking Jobs  | Stark  | Data Platform Engineer – Data...

  • About Us

    STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments.

    We're focused on delivering deployable, high-performance systems — not future promises. In a time of rising threats, STARK is bolstering the technological edge of NATO Allies and their Partners to deter aggression and defend Europe — today.


    About the team

    The Data Operations team owns the entire data lifecycle behind STARK's AI stack: collection, acquisition, generation, curation, and management. We run our own data-collection campaigns across Europe, evaluate new sensors and platforms, and build the internal data platform that turns raw recordings into ready-to-use datasets. Everything we produce feeds directly into the perception and autonomy systems deployed on STARK's platforms — a real data advantage is built, not bought. The team is scaling up right now: real scope, direct impact, no legacy.


    Your mission

    Data is the fuel of STARK’s AI stack — you build the engine that makes it usable. You own the software backbone of our data platform: the metadata systems, ETL pipelines, data contracts, catalogs, databases, and internal tools that let engineers find, understand, validate, and reuse terabytes of multi-sensor field data in minutes, not days. You treat data context as a product: structured, searchable, version-aware, documented, and traceable from raw recording to processed asset, annotation delivery, dataset, and downstream ML workflow. Today, much of this is manual, scattered, or implicit — your job is to automate it away, support labeling efforts with the right data tooling, and turn operational data into reliable systems.


    Responsibilities

    • Design, implement, and maintain our metadata database and data catalog (datasets, recordings, sensors, labels, lineage)

    • Build and operate ETL/ingest pipelines that bring field recordings, synthetic data, and external deliveries into our cloud storage (GCP)

    • Own the data management and labeling lifecycle end-to-end: coordinate and communicate with external labeling companies and data subcontractors, track deliveries, run QA reports, and build the operational workflows they work in

    • Develop internal enabling tools for the whole AI organization: dataset search and filtering, APIs/backend, dashboards, and self-service data access

    • Run data migrations and indexing jobs; keep the catalog consistent and fast as data volume grows

    • Handle admin support and user access management — and then automate these support tasks so they stop being manual work

    • Establish good engineering hygiene in a young codebase: tests, typing, docs, logging, CI/CD

    • Shape the long-term architecture and vision of the data platform together with the team


    Qualifications

    • Strong Python

    • Solid SQL/PostgreSQL, including schema design

    • Experience with data modeling and metadata systems

    • Experience designing and operating ETL/data pipelines

    • Docker and CI/CD basics

    • Hands-on with object storage (GCS, S3, or similar)

    • Good software engineering hygiene: tests, docs, typing, logging

    • Organized and pragmatic: you can prioritize between a quick fix and a proper solution, and you know when each is right

    • Not allergic to support tasks — but technical enough to automate the support away

    • Comfortable coordinating with external vendors and non-technical stakeholders


    Nice to have

    • Familiarity with ML datasets and labeling workflows (images, video, lidar; annotation formats like COCO)

    • Experience with synthetic data generation or GenAI-assisted data workflows (auto-labeling, data augmentation, foundation-model-based curation)

    • Experience with GCP services beyond storage (BigQuery, Cloud Run, IAM)

    • Experience with data versioning / dataset tooling (DVC, LakeFS, FiftyOne, or similar)

    • Experience in a startup environment — comfortable with ambiguity and changing priorities

    • Exposure to robotics data formats (ROS bags, MCAP, PX4 logs)

    Jobs at Stark

    All Jobs at Stark →

    Job recommendations