Lead Analytics Engineer/Data Analyst – AI & Foundation Models
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Lead Analytics Engineer/Data Analyst – AI & Foundation ModelsOverviewMastercard is seeking a Lead Analytics Engineer / Data Analyst to join an AI product team building cutting-edge foundation model capabilities with enterprise-wide impact. This role bridges the gap between the organization's governed, enterprise-modelled data and the AI engineers and data scientists who need to use data confidently for model training and evaluation. Today, this data is often accessed without a clear, standard path — this role will design and own the curated, documented layer that changes that, making the right data easy to find, understand, trust, and use correctly. You will join a team of AI Engineers, Software Developers, Data Scientists and MLOps to bring these capabilities to market.
Role
In this role, you will be responsible for turning already-governed enterprise data into a standardized, trustworthy, and well-documented resource for AI model training and evaluation. The role is less focused on owning database/ETL infrastructure or technologies and more focused on ensuring our AI team is utilizing enterprise data effectively.
Key responsibilities:
• Design, build, and maintain a curated, standardized layer of training-ready datasets, so AI engineers and data scientists have a clear, consistent, well-understood path to the data they need. Replacing ad hoc, unguided data selection from our data sources.
• Create final stages of ETL pipelines that utilize gold, silver and bronze layers to provide data as a service to our Data Scientists within Databricks.
• Develop deep expertise in the organization's existing modelled data. Schemas, meaning, lineage, and quality. Sufficient to guide others confidently and design genuinely useful curated assets.
• Partner directly with data scientists and AI engineers to understand recurring training and model development data needs, and convert ad hoc, repeated requests into reusable, documented, standard data processes.
• Own documentation and cataloging for these curated assets, ensuring they are genuinely discoverable and understandable, not just technically accessible.
• Assemble and shape specific training data extracts. Joins, filters, and dataset-level transformations. For use cases not yet covered by an existing standard asset.
• Build insight and reporting views (e.g., model monitoring, cost tracking, evaluation results) using Databricks (preferred) or Microsoft Fabric.
• Act as the primary point of contact for AI and data science teams asking "where do I find this data" and "can I trust this dataset."
• Contribute to the long-term evolution of this capability toward more formalized data/feature product practices as the platform and its use cases mature.
• Coordinate with the enterprise data engineering team that owns upstream data modelling and pipelines. Representing the AI platform's data needs and surfacing quality or schema issues as they arise.
• Support data-related troubleshooting when model training or evaluation issues trace back to a data problem.
All About You:
• Expert-level SQL
• Strong data modelling and semantic-layer instincts. Able to design clear, sensible, well-structured curated views and datasets, even without owning the underlying raw-to-gold transformation.
• Genuine skill at navigating and documenting a large, complex data environment built by others. Reading schemas, tracing lineage, and understanding data you didn't create.
• Unity Catalog fluency. Metadata, tagging, documentation, lineage, and access patterns.
• Strong communication and translation skills. Comfortable working directly with technical AI/data science stakeholders to turn an ambiguous data need into a concrete, well-designed, reusable asset.
• Experience building reporting/dashboards, ideally within Databricks; Power BI/Microsoft Fabric experience is a strong and acceptable alternative.
• Working understanding of what makes data "training-ready" for ML use cases: e.g., feature completeness, avoiding data leakage, temporal correctness. Even without building or training models directly.
• Strong documentation discipline. Treats clear documentation as a core deliverable, not an afterthought.
• Comfortable operating without infrastructure ownership. Thrives working within centrally-owned data pipelines and platforms rather than building or maintaining them.
• A genuine interest in growing this capability over time, as it evolves from curated datasets toward more formalized data and feature product practices.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Abide by Mastercard’s security policies and practices;
Ensure the confidentiality and integrity of the information being accessed;
Report any suspected information security violation or breach, and
Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.