Point your AI agent at freehire and let it find you a job.

Get the CLI →

Jobgether

NewBe an early applicant

Lead Data Engineer

Posted
Discussion

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Data Engineer based in India.

As a Lead Data Engineer, you’ll architect and own the end-to-end data sourcing and ETL infrastructure powering a large-scale B2B data platform.
You’ll transform complex, unstructured web data into accurate, structured datasets across multiple international markets.
The role combines advanced data engineering, web extraction, LLM infrastructure, agentic workflows, and production-grade orchestration.
You’ll establish the technical standards and frameworks that enable other engineers to build reliable, scalable extraction systems.
A significant part of the role is hands-on, tackling difficult problems such as anti-bot defenses, JavaScript-heavy sources, schema drift, and cost optimization.
You’ll also shape how AI is evaluated and operated in production through versioning, observability, labelled datasets, and precision and recall measurement.
Working within a lean, senior, AI-native environment, you’ll have substantial autonomy and direct ownership of architecture, quality, performance, and outcomes.

Accountabilities

  • Architect and own the complete data sourcing and ETL layer, covering extraction, normalization, deduplication, validation, loading, reliability, performance, and cost.

  • Define and maintain standards for data coverage, accuracy, schema design, and the appropriate use of deterministic rules versus LLM-based processing.

  • Design scalable extractor frameworks, abstractions, patterns, and tooling that enable engineers to rapidly build reliable new extraction systems.

  • Develop and ship complex extractors capable of handling anti-bot defenses, proxy and IP rotation, JavaScript-heavy rendering, schema drift, and poorly structured sources.

  • Build agentic extraction workflows incorporating rule-based triage, LLM escalation, structured-output validation, retries, and human-review queues.

  • Establish production-grade LLM infrastructure, including versioned prompts, labelled evaluation datasets, precision and recall measurement, traceability, rollback mechanisms, and model-drift detection.

  • Develop evaluation and observability frameworks that make LLM-powered extraction measurable, debuggable, and reliable in production.

  • Establish a model-selection strategy that balances extraction quality, latency, and cost according to the complexity of each task.

  • Build and maintain a versioned skills library containing reusable extraction specifications and patterns that can be used by both engineers and AI agents.

  • Design and operate scalable orchestration using Airflow or equivalent technologies, including dependency management, recovery mechanisms, incremental processing, scheduling, and cost controls.

  • Provide technical input to data platform, product, and analytics teams on sourcing architecture, coverage, schema, and data accuracy.

  • Maintain a strong focus on customer-facing data quality, making informed decisions about when to refactor, optimize, ship, or redesign systems.

  • Contribute hands-on reference implementations and solve the most technically challenging extraction problems while establishing patterns for the wider engineering team.

  • Requirements

    • 6–10 years of experience building and shipping production extraction, ETL, or data pipeline systems in business-critical environments.

    • Deep experience with large-scale web data extraction, including anti-bot defenses, proxy architectures, JavaScript rendering, schema drift, recovery strategies, and complex source structures.

    • Demonstrated experience deploying LLMs as production components within extraction pipelines, including structured outputs, versioned prompts, labelled evaluation sets, traces, precision/recall measurement, and model-drift detection.

    • Proven experience designing and operating agentic workflows and tool-calling systems, including reusable skill specifications and production implementations.

    • Strong architectural judgment regarding when to use deterministic rules versus LLM-based approaches, with the ability to establish and communicate these standards across a team.

    • Expert-level Python skills, including production-grade development, concurrency, performance optimization, and large-scale processing.

    • Advanced SQL skills, particularly for complex transformations, analytical workloads, and performance optimization.

    • Extensive experience with Airflow or an equivalent orchestration platform, including dependencies, recovery patterns, scheduling, and production-scale cost management.

    • Hands-on experience using AI-native development tools such as Claude Code, Cursor, or equivalent platforms, with production systems or implementations to demonstrate.

    • Strong systems architecture and product judgment, with the ability to balance technical quality, delivery speed, data accuracy, customer impact, and operational cost.

    • A practical, hands-on approach to AI tooling, structured outputs, evaluations, tracing, observability, and LLM-powered infrastructure.

    • Experience with cloud data platforms such as Snowflake or Redshift and AWS services including Lambda, S3, ECS, or Glue is highly valued.

    • Familiarity with vector databases, embeddings, retrieval systems, matching, deduplication, automated data-quality frameworks, or anomaly detection is advantageous.

    • Experience with LLM evaluation and tracing tools such as Braintrust, Promptfoo, Inspect, Logfire, OpenTelemetry, or comparable technologies is beneficial.

    • Familiarity with B2B data, including firmographics, people data, and company registries across multiple markets, is a plus.

    • Understanding of data privacy and compliance frameworks such as GDPR and CCPA is advantageous.

    • Startup or scaleup experience, particularly in environments where engineers own outcomes end to end and ship rapidly, is highly valued.

    • Comfortable working autonomously within a lean, senior team and making architecture decisions without extensive process or predefined playbooks.

    • Benefits

      • Fully remote position based in India.

      • High level of autonomy over technical architecture and implementation decisions.

      • Opportunity to architect a largely greenfield AI-native data sourcing platform from first principles.

      • Hands-on exposure to advanced LLM infrastructure, agentic extraction, evaluation systems, and production AI workflows.

      • Opportunity to solve technically challenging data extraction problems across multiple international markets.

      • Significant ownership and direct impact on the quality, accuracy, speed, and cost of a core data platform.

      • Lean, senior-team environment with minimal organizational layers and a strong focus on rapid delivery.

      • Flexible working approach with no fixed working hours.

      • Opportunity to establish technical standards and reference implementations that other engineers build upon.

      • Strong career growth potential through ownership of a critical engineering function and measurable business outcomes.

      • Opportunity to work in an AI-native environment where modern AI development and observability practices are embedded into everyday engineering.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Skills

What Lead Data Engineering jobs ask for — and how much of it you have →

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available