32_【海外人材枠】Data Engineer ー AI-Era Corporate Data Infrastructure | Full Remote(Vietnam)
Posted Updated
株式会社SalesNow 32_【海外人材枠】Data Engineer ー AI-Era Corporate Data Infrastructure | Full Remote(Vietnam)
Data Engineer — No Japanese required, 100% remote for a Japanese AI company that builds its own product|An 8 billion record data platform|Vietnamese engineers already on the team, hiring 5-10 more
仕事概要
## <About SalesNow>SalesNow's mission is 「誰もが活躍できる仕組みをつくる。」 (Create systems where everyone can thrive), and we are taking on the challenge of fundamentally changing how people work. Japan's workforce is shrinking, so the question is how much more value one person can create. SalesNow answers it by rebuilding how business-to-business work is done, using corporate data and AI.
The foundation is a database of over 16 million corporate and organizational records, uniquely structured down to companies, organizations, branches, and departments. The data infrastructure behind it holds 8 billion records. SalesNow ranks No.1 in corporate database record count (企業データベース収録件数No.1; period ending October 2025; research by Japan Marketing Research Organization / 日本マーケティングリサーチ機構).
## <Why This Role Exists>
**The era when revenue grew in proportion to headcount is over.** More of the value one person creates now comes from working together with AI, and AI can only support good decisions when it works on accurate data about each company. Building that data is the job.
A large part of this data never appears on the web at all: it comes from offline collection by tens of thousands of researchers and our own in-house team, and from contributions by users of our own media. Resolving a single company identity out of sources like these, rather than out of a public dataset, is the engineering problem you take on here. SalesNow is approaching Series B, and we have set a goal to grow ARR per employee tenfold in three years, starting from November 2025. That goal only works if the data is right, which puts data engineering at the center of how this company grows. We are hiring data engineers who design how corporate data is collected, integrated, and verified, and who make it usable by both people and AI.
### The team in Vietnam
Three Vietnamese members already work with us, and we plan to hire five to ten more. You join a cross-language workflow that already runs, and you help shape how the Vietnam team works as it grows.
## <What You Will Build>
- Company profiles that combine company overview, job postings, IR, services, technologies and tools in use, advertising, organizations and departments, and contact details, matched down to branch level with our own JC code. You decide how these sources merge so each company ends up with exactly one profile
- Signals that capture company moves, such as hiring increases, funding rounds, new offices, overseas expansion, downward earnings revisions, and paused job postings. Sales teams use these signals to decide which companies to approach and when. You design the validation checks that make those signals reliable enough to act on
- Airflow (MWAA) DAGs for our job-posting sources, where you decide how anomalies are detected, how runs recover, and what "collected correctly" means for each source
- A Databricks and Delta Lake platform in three layers (raw, cleansed, serving), with job definitions in Terraform. Where a record is cleaned, and what is allowed into the serving layer, are your design calls
- Data delivered through the SalesNow web app, Salesforce and HubSpot integrations, the Data API, SalesNow MCP (used directly from AI agents such as Claude and Cursor), and AI implementation projects for customers. A schema decision you make shows up both in a customer's CRM and in the answers an AI agent gives
## <Your Technical Challenges>
### Data collection and ingestion
- Design and implement large-scale web scraping systems that reliably extract structured data from many heterogeneous sources
- Build API integration pipelines for partner data feeds, with schema validation and anomaly detection
- Architect real-time ingestion pipelines for daily updates with sub-minute latency targets
### Data transformation and quality
- Build and maintain dbt models that transform raw data into the unified SalesNow schema
- Define data quality SLAs (accuracy, completeness, freshness) and build monitoring dashboards that alert before customers notice
- Implement entity resolution and deduplication at scale: matching company records across sources with different formats, naming conventions, and identifiers
### Infrastructure and orchestration
- Orchestrate complex Airflow DAGs across collection, transformation, enrichment, and delivery
- Optimize PostgreSQL and Elasticsearch / OpenSearch clusters for 8 billion records with millisecond query response
- Design fault-tolerant pipelines that recover automatically from failures, with clear observability at every stage
### AI pipeline integration
- Feed structured data into RAG pipelines and AI agent systems via SalesNow MCP
- Build data delivery layers that AI agents query in real time for enterprise customer workflows
- Work with Amazon Bedrock, Claude, and Gemini to improve how AI systems use structured corporate data
### Problems we are working on now
- Decide whether a job posting is still open. Closure-detection logic is split across three systems, and behavior differs by source. Some sources do not show closure status on their listing pages, so closing postings by diff can wrongly close them when a crawl misses a page. For each source, you weigh the risk of marking an active posting as closed against the risk of leaving an expired posting open
- Design the process that collects every company's website information and keeps tracking changes, so company overviews and service information stay current
### Verifying our own data against evidence
When we had AI independently verify our logic for inferring the client company behind freelance job postings, only 12 of 100 cases could be identified with confirmed evidence. We revised the logic so that cases without enough evidence are marked "cannot identify." Here you are expected to question the data, including the output of your own pipelines, and when the evidence falls short, you have the authority to change the logic.
### Scope of duties
Scope of changes to the duties you will perform: yes (従事すべき業務の変更の範囲:有り)
## <How We Work with AI>
- Everyone can use Claude and Gemini, and we also use Codex. You choose the model that fits each task
- We ask multiple LLMs to review our work, and people check the evidence and decide what to adopt. For data checks, we use LLM as a Judge with criteria that people set. CodeRabbit runs a first review on pull requests
- Instructions written for AI go through pull request review, the same as code, so prompts and agents are engineered with the same rigor as production pipelines. Deliverables are managed on GitHub
- Once the goal and budget are agreed, the person in charge decides how to proceed, where to use AI, and which tools to use. For data collection, you also own what each record costs to obtain, and every pipeline you design is judged on both accuracy and unit cost
- Every role is expected to redesign its own work with AI. Our leadership team writes code too, so you discuss technical trade-offs with people who build
- Across the company, 2,229 pull requests were merged in the 30 days up to 12 September 2026. At that pace, you judge AI-assisted pipeline changes against real data, failure modes, and operating costs
## <Tech Stack>
- Languages: Python (FastAPI), SQL, Scala
- Data transformation: dbt
- Orchestration: Airflow (MWAA)
- Data processing: Databricks (Spark / Delta Lake), Scrapy
- Infrastructure as code: Terraform
- Databases: PostgreSQL, Elasticsearch, OpenSearch
- Cloud: AWS (primary), Vercel
- AI/ML: Amazon Bedrock, Claude, Gemini
- Development tools: Claude Code, Codex, Cursor, GitHub Copilot, CodeRabbit
- CI/CD: GitHub Actions
- Communication: Slack, GitHub, Notion
## <Working Across Languages>
- No Japanese required. English is the working language for engineering: code, pull requests, technical documents, and standups
- Work comes with a PRD. Both sides agree on the requirements and the deadline before you start, so you spend your time on design and code
- GitHub, Notion, and Slack are translated. Pull requests, documents, and daily discussions all pass through translation, so you follow every discussion without Japanese
- Company-paid AI translation. Members who do not speak Japanese discuss work with Japanese colleagues through AI simultaneous interpretation, and interviews run the same way
## <Hiring Process>
1. Application review (resume + brief answers to 3 technical questions)
2. One-day selection event, completed in a single day: a data pipeline exercise on a realistic scenario (2-3 hours), a technical interview (60 min) covering your past work and system design, and a culture fit conversation with the engineering lead (30 min)
3. Offer
Timeline: 2-3 weeks from application to offer
## <Equal Opportunity>
We are an equal opportunity employer and evaluate candidates on skills, experience, and potential, regardless of nationality, gender, age, or background.
## <Links>
- Interview with an engineer on our team, who joined as an intern and now develops SalesNow MCP and our API (Japanese): https://note.com/salesnow/n/na813feacc9cb
- Culture deck (Japanese): https://speakerdeck.com/salesnow/culture-deck
必須スキル
## <About SalesNow>SalesNow's mission is 「誰もが活躍できる仕組みをつくる。」 (Create systems where everyone can thrive), and we are taking on the challenge of fundamentally changing how people work. Japan's workforce is shrinking, so the question is how much more value one person can create. SalesNow answers it by rebuilding how business-to-business work is done, using corporate data and AI.
The foundation is a database of over 16 million corporate and organizational records, uniquely structured down to companies, organizations, branches, and departments. The data infrastructure behind it holds 8 billion records. SalesNow ranks No.1 in corporate database record count (企業データベース収録件数No.1; period ending October 2025; research by Japan Marketing Research Organization / 日本マーケティングリサーチ機構).
## <Why This Role Exists>
**The era when revenue grew in proportion to headcount is over.** More of the value one person creates now comes from working together with AI, and AI can only support good decisions when it works on accurate data about each company. Building that data is the job.
A large part of this data never appears on the web at all: it comes from offline collection by tens of thousands of researchers and our own in-house team, and from contributions by users of our own media. Resolving a single company identity out of sources like these, rather than out of a public dataset, is the engineering problem you take on here. SalesNow is approaching Series B, and we have set a goal to grow ARR per employee tenfold in three years, starting from November 2025. That goal only works if the data is right, which puts data engineering at the center of how this company grows. We are hiring data engineers who design how corporate data is collected, integrated, and verified, and who make it usable by both people and AI.
### The team in Vietnam
Three Vietnamese members already work with us, and we plan to hire five to ten more. You join a cross-language workflow that already runs, and you help shape how the Vietnam team works as it grows.
## <What You Will Build>
- Company profiles that combine company overview, job postings, IR, services, technologies and tools in use, advertising, organizations and departments, and contact details, matched down to branch level with our own JC code. You decide how these sources merge so each company ends up with exactly one profile
- Signals that capture company moves, such as hiring increases, funding rounds, new offices, overseas expansion, downward earnings revisions, and paused job postings. Sales teams use these signals to decide which companies to approach and when. You design the validation checks that make those signals reliable enough to act on
- Airflow (MWAA) DAGs for our job-posting sources, where you decide how anomalies are detected, how runs recover, and what "collected correctly" means for each source
- A Databricks and Delta Lake platform in three layers (raw, cleansed, serving), with job definitions in Terraform. Where a record is cleaned, and what is allowed into the serving layer, are your design calls
- Data delivered through the SalesNow web app, Salesforce and HubSpot integrations, the Data API, SalesNow MCP (used directly from AI agents such as Claude and Cursor), and AI implementation projects for customers. A schema decision you make shows up both in a customer's CRM and in the answers an AI agent gives
## <Your Technical Challenges>
### Data collection and ingestion
- Design and implement large-scale web scraping systems that reliably extract structured data from many heterogeneous sources
- Build API integration pipelines for partner data feeds, with schema validation and anomaly detection
- Architect real-time ingestion pipelines for daily updates with sub-minute latency targets
### Data transformation and quality
- Build and maintain dbt models that transform raw data into the unified SalesNow schema
- Define data quality SLAs (accuracy, completeness, freshness) and build monitoring dashboards that alert before customers notice
- Implement entity resolution and deduplication at scale: matching company records across sources with different formats, naming conventions, and identifiers
### Infrastructure and orchestration
- Orchestrate complex Airflow DAGs across collection, transformation, enrichment, and delivery
- Optimize PostgreSQL and Elasticsearch / OpenSearch clusters for 8 billion records with millisecond query response
- Design fault-tolerant pipelines that recover automatically from failures, with clear observability at every stage
### AI pipeline integration
- Feed structured data into RAG pipelines and AI agent systems via SalesNow MCP
- Build data delivery layers that AI agents query in real time for enterprise customer workflows
- Work with Amazon Bedrock, Claude, and Gemini to improve how AI systems use structured corporate data
### Problems we are working on now
- Decide whether a job posting is still open. Closure-detection logic is split across three systems, and behavior differs by source. Some sources do not show closure status on their listing pages, so closing postings by diff can wrongly close them when a crawl misses a page. For each source, you weigh the risk of marking an active posting as closed against the risk of leaving an expired posting open
- Design the process that collects every company's website information and keeps tracking changes, so company overviews and service information stay current
### Verifying our own data against evidence
When we had AI independently verify our logic for inferring the client company behind freelance job postings, only 12 of 100 cases could be identified with confirmed evidence. We revised the logic so that cases without enough evidence are marked "cannot identify." Here you are expected to question the data, including the output of your own pipelines, and when the evidence falls short, you have the authority to change the logic.
### Scope of duties
Scope of changes to the duties you will perform: yes (従事すべき業務の変更の範囲:有り)
## <How We Work with AI>
- Everyone can use Claude and Gemini, and we also use Codex. You choose the model that fits each task
- We ask multiple LLMs to review our work, and people check the evidence and decide what to adopt. For data checks, we use LLM as a Judge with criteria that people set. CodeRabbit runs a first review on pull requests
- Instructions written for AI go through pull request review, the same as code, so prompts and agents are engineered with the same rigor as production pipelines. Deliverables are managed on GitHub
- Once the goal and budget are agreed, the person in charge decides how to proceed, where to use AI, and which tools to use. For data collection, you also own what each record costs to obtain, and every pipeline you design is judged on both accuracy and unit cost
- Every role is expected to redesign its own work with AI. Our leadership team writes code too, so you discuss technical trade-offs with people who build
- Across the company, 2,229 pull requests were merged in the 30 days up to 12 September 2026. At that pace, you judge AI-assisted pipeline changes against real data, failure modes, and operating costs
## <Tech Stack>
- Languages: Python (FastAPI), SQL, Scala
- Data transformation: dbt
- Orchestration: Airflow (MWAA)
- Data processing: Databricks (Spark / Delta Lake), Scrapy
- Infrastructure as code: Terraform
- Databases: PostgreSQL, Elasticsearch, OpenSearch
- Cloud: AWS (primary), Vercel
- AI/ML: Amazon Bedrock, Claude, Gemini
- Development tools: Claude Code, Codex, Cursor, GitHub Copilot, CodeRabbit
- CI/CD: GitHub Actions
- Communication: Slack, GitHub, Notion
## <Working Across Languages>
- No Japanese required. English is the working language for engineering: code, pull requests, technical documents, and standups
- Work comes with a PRD. Both sides agree on the requirements and the deadline before you start, so you spend your time on design and code
- GitHub, Notion, and Slack are translated. Pull requests, documents, and daily discussions all pass through translation, so you follow every discussion without Japanese
- Company-paid AI translation. Members who do not speak Japanese discuss work with Japanese colleagues through AI simultaneous interpretation, and interviews run the same way
## <Hiring Process>
1. Application review (resume + brief answers to 3 technical questions)
2. One-day selection event, completed in a single day: a data pipeline exercise on a realistic scenario (2-3 hours), a technical interview (60 min) covering your past work and system design, and a culture fit conversation with the engineering lead (30 min)
3. Offer
Timeline: 2-3 weeks from application to offer
## <Equal Opportunity>
We are an equal opportunity employer and evaluate candidates on skills, experience, and potential, regardless of nationality, gender, age, or background.
## <Links>
- Interview with an engineer on our team, who joined as an intern and now develops SalesNow MCP and our API (Japanese): https://note.com/salesnow/n/na813feacc9cb
- Culture deck (Japanese): https://speakerdeck.com/salesnow/culture-deck
歓迎スキル
- Experience with Elasticsearch or OpenSearch at scale- dbt experience
- Web scraping or large-scale data collection systems
- Experience with data quality frameworks and monitoring
- AWS experience (S3, Lambda, ECS, RDS, etc.)
- RAG pipeline or vector database experience
- Experience processing 1M+ records per day
- Japanese language ability (JLPT N3 or above). Not required, but it lets you work more closely with business teams and widens your career options within the company
求める人物像
- You share SalesNow's mission 「誰もが活躍できる仕組みをつくる。」 (Create systems where everyone can thrive) and our values (AI Native, Kotoshikou (purpose-driven), and Highest Quality), and you are determined to make full use of AI- You enjoy moving fast. Even as things change, you start from the purpose, find the problem, and see it through to a solution
- You work with your team and customers with integrity and ownership
- You care about data quality and protect it with systems, not manual checks alone
- You verify AI output against evidence before adopting it
- You can explain your design decisions, including their cost, to teammates across languages
**In two to three years, using AI will be a given. The difference will come from whether you can design work and organizations with AI as the starting point. At SalesNow, that is what we do every day.**
応募概要
給与
USD 1,000 – 2,500 per month (25 – 62 million VND per month), based on experience and skills勤務地
100% remote. Work from Ho Chi Minh City, Hanoi, Da Nang, or anywhere with a stable internet connection雇用形態
Full-time, remote contractor via EOR (Employer of Record) — full legal employment with local benefits勤務体系
- Work hours: 10:00-19:00 Japan time, which is 8:00-17:00 Vietnam time- Public holidays: You take the Vietnamese public holidays (Tet, Hung Kings' Festival, 30 April, 1 May, 2 September), not the Japanese calendar
- Paid leave: Per Vietnamese labor law
- Career path: Performance-based salary reviews every 6 months. Path to Senior DE, Lead DE, and Data Architect
- Work style: Fully remote
- Scope of changes to work location: none (就業場所の変更の範囲:無し)
試用期間
福利厚生
- AI tools: spending averages more than USD 600 per person per month (company-paid)- Hardware: MacBook provided
Skills
- Agentic AI
- AI
- Airflow
- Anomaly Detection
- API
- AWS
- AWS Bedrock
- CI/CD
- Claude Code
- Cloud
- CRM
- Data Engineering
- Data Pipelines
- Data Quality
- Databricks
- dbt
- Delta Lake
- ECS
- Elasticsearch
- FastAPI
- GitHub
- GitHub Actions
- Github Copilot
- HubSpot
- Infrastructure as Code
- Lambda
- LLM
- Machine Learning
- MCP
- Notion
- Observability
- OpenSearch
- PostgreSQL
- Python
- RDS
- S3
- Salesforce
- Scala
- Slack
- Spark
- SQL
- Terraform
- Vector Databases