Data Engineer
Posted Updated
Remote position - only for professionals based in LATAM.
We are looking for a Senior Data Engineer for one of our client's teams. You will take ownership of the systems responsible for ingesting, transforming, validating, and publishing data across our platform.
In this role, you will work at the data ingestion boundary, where data from multiple external sources enters our systems. You will be responsible for building reliable pipelines, resolving identity and entity-matching challenges, detecting data-quality issues, and ensuring data flows correctly into downstream products and services.
This is a hands-on engineering role that combines AWS data engineering, Python/SQL development, event-driven architectures, data quality, and identity resolution. You will collaborate closely with engineering, product, analytics, and downstream platform teams to troubleshoot issues and improve the reliability of our data ecosystem.
What You'll Do
• Own and evolve data ingestion pipelines, including source ingestion/scraping, data cleaning and curation using AWS Glue, entity resolution, and event-driven data publishing.
• Design and maintain reliable batch and event-driven data workflows across our data platform.
• Diagnose and resolve identity-resolution issues across multiple data sources, including deduplication and entity-matching challenges.
• Develop and maintain data quality and validation frameworks, including freshness monitoring, null-rate checks, consistency validation, and schema-drift detection.
• Monitor data at the ingestion boundary and proactively identify issues before they impact downstream systems and products.
• Partner with engineering teams responsible for the events bridge and downstream identity services to trace data and events end-to-end.
• Collaborate with Product, Assessments, Analytics, and other stakeholders to understand how data flows into downstream products and ensure those requirements are reflected in the data pipelines.
• Automate infrastructure and pipeline changes using Infrastructure as Code, primarily Terraform or AWS CDK.
• Work with AWS services including Glue, Athena, S3, and event-driven services to build and operate scalable data infrastructure.
• Participate in production incident response, quickly identifying the scope and impact of data-quality issues and implementing remediation.
• Continuously improve pipeline reliability, observability, maintainability, and operational processes.
What You'll Bring
• 5+ years of experience building and operating production-grade data pipelines.
• Strong experience working with both batch data processing and event-driven architectures.
• Hands-on experience with AWS Glue, Athena, and S3-based data lakes, including layered/medallion-style data transformations.
• Strong proficiency in Python and SQL for data transformation, processing, and analysis.
• Experience with identity resolution, entity matching, or deduplication, including exact and fuzzy matching approaches.
• Understanding of the challenges and tradeoffs involved in maintaining durable identifiers and first-seen/locked identity mappings.
• Experience working with event schemas and schema-registry-backed contracts, such as Protobuf, and an understanding of the risks associated with schema changes.
• Strong troubleshooting and production incident-response skills, including the ability to determine impact, identify root causes, and implement fixes quickly.
• Ability to understand and debug code written in a functional or concurrent programming language, such as Elixir.
• Strong communication and collaboration skills, with the ability to work effectively across engineering, product, analytics, and other technical teams.
Nice to Have
• Experience with Elixir/Phoenix or another BEAM-based concurrent processing framework.
• Experience with DynamoDB-backed identity, lookup, or matching services.
• Experience with the Snowflake ecosystem, including data modeling, Snowpipe, Streams, and Tasks.
• Experience working with sports data providers, such as or Sportradar, or with other licensed-content/data provider ecosystems.
• Experience implementing data observability, including freshness/staleness alerts, null-rate monitoring, schema-drift detection, and data-quality dashboards.
• Experience working with Infrastructure as Code using Terraform or AWS CDK.
Technologies
Languages: Python, SQL, familiarity with Elixir
AWS: Glue, Athena, S3, DynamoDB
Data & Streaming: Data Lakes, Event-Driven Architecture, Protobuf, Schema Registries
Infrastructure: Terraform, AWS CDK
Data Platforms: Snowflake
Engineering Practices: Data Quality, Data Observability, Identity Resolution, Entity Matching, Incident Response