27_【海外人材枠/正社員】データエンジニア(Data Engineer)
Posted Updated
Data Engineer (Open to Global Talent) — Generative AI × Databricks for High-Quality Data Infrastructure | Over 16 Million Corporate and Organizational Records | Full Remote, Full Flex
仕事概要
This position is open to non-Japanese candidates who can work in Japan as full-time employees and use Japanese at a business level.## <About SalesNow>
SalesNow's mission is 「誰もが活躍できる仕組みをつくる。」 (Create systems where everyone can thrive), and we are taking on the challenge of fundamentally changing how people work. We rebuild how B2B work gets done, using corporate data and AI so that each person can create more value. SalesNow is the corporate data infrastructure for the AI era, supplying the data that people and AI agents work from. Our database covers over 16 million corporate and organizational profiles, structured down to individual offices and departments, and built on 8 billion records. SalesNow ranks No.1 in corporate database record count (企業データベース収録件数No.1; survey period: October 2025; research by Japan Marketing Research Organization).
SalesNow Data Network, our own collection network, combines three routes. Our platform uses AI to update web information daily. A partner network of tens of thousands of data researchers, working alongside our own in-house team, collects offline — from on-site surveys to reading paper documents. Our own corporate-information media draws 7 million page views a month, and that user network also takes on part of the offline collection. Together they give us what search engines never return, including org charts and more than 7.5 million department contacts.
Our own JC code assigns a unique ID to each site — headquarters, plants, branches, and stores — not only to the company. Multi-site groups can therefore be traced accurately. Records are matched on corporate number, name, address, phone, URL, and email. Fifteen levels of priority resolve them to a single company. For unresolved company matches, our team uses an AI workflow to supplement company information from current web sources. The platform can attach more than 50 attributes to company records, including registry information, industry classification, revenue, capital, headcount with its rate of change, sites, contacts, and scores. This collection, matching, and validation work is what you will design at production scale.
Sales teams use signals such as a rise in hiring, a funding round, a new office, or a stop in job postings to see why a company is worth approaching now. Customers use the data in the corporate database SalesNow, in their Salesforce and HubSpot records, and from AI agents such as Claude and Cursor through SalesNow MCP. Our Data API (beta) ships as four APIs: Company, News, Recruit, and Organization (departments, org charts, and sites). Companies that have adopted SalesNow include Yamato Transport, Panasonic, LY Corporation, SMBC Nikko Securities, PERSOL CAREER, JCB, and GMO Payment Gateway. You develop your data engineering expertise on a product with enterprise adoption, where the matching and validation you design decide what those customers can rely on.
## <Why We Are Hiring>
**The era when revenue grew in proportion to headcount is over.** As Japan's workforce shrinks, each person needs to create more value. We have set a goal to grow ARR per employee tenfold in three years, starting from November 2025. On the data side, that means redesigning how corporate data is collected, matched, and verified, rather than adding people to check it.
SalesNow is approaching Series B. You will set the collection requirements and the quality criteria behind SalesNow MCP and the Data API: which sources to trust, what evidence a record needs, and when a pipeline result is good enough to ship.
The same record now serves the SalesNow screen, customers' Salesforce and HubSpot data, the Data API, and AI agents through SalesNow MCP. We need an engineer who owns collection, matching, and verification end to end, because a quality rule set at one stage decides what all four of those consumers receive.
## <Mission>
You will build structured company profiles that capture what each company does and what has changed recently. You will tell apart companies with the same name and handle address variations, mergers, and relocations. You define the matching rules that distinguish separate companies and link records belonging to the same company. Customers rely on those records when they research companies in SalesNow, or through AI agents connected via SalesNow MCP.
## <What You Will Do>
You will build and improve our corporate data platform, which runs at the scale of tens of terabytes. You will take on one or more of the following:
- Build and improve pipelines for web data crawling
- Implement data cleansing and enrichment
- Build a monitoring environment to improve data quality
- Improve the data platform for scalability, performance, and cost efficiency
- Design the data platform and data quality foundations with future data science use in mind
### Examples of problems you will work on
- Decide whether a job posting is still open. Closure detection is currently split across three systems and behaves differently by source. Some sources do not show closures on their listing pages. If disappearance from a crawl triggers closure, a single missed page can incorrectly mark a posting as closed. You will set criteria for each source that balance false closure detections against missed closures
- Keep company website information current. You will design the process that fetches information from company websites and keeps tracking what has changed since the last fetch, so that each company's profile stays up to date
※Scope of changes to duties: yes (従事すべき業務の変更の範囲:有り)
## <Tech Stack>
- Data collection: Python (Scrapy), AWS Fargate for Amazon ECS, Amazon Managed Workflows for Apache Airflow
- Data storage: Amazon S3, Amazon Aurora PostgreSQL
- Data transformation: Scala and Python on Databricks (Apache Spark, Delta Lake; processed in three layers: raw, cleansed, and serving)
- Infrastructure: AWS, Databricks
- Other: GitHub, Slack, Notion, Asana, Google Workspace
- Development tools: Claude Code, Codex, Cursor, CodeRabbit, n8n
## <How We Work>
### Working with AI
- You use company-funded AI tools to build and verify pipelines; spending averages over JPY 100,000 per person per month
- Everyone can use Claude and Gemini. We also use Codex. You compare models on matching and extraction tasks, and select them based on your own quality checks
- You have several LLMs review your pipeline code and data rules, and you decide which findings to adopt after checking the evidence yourself
- Engineers use CodeRabbit for first-pass PR reviews, and focus their own review on design validity, data assumptions, impact, edge cases, data quality, and reproducibility
- In data processing, we structure listed companies' disclosure documents with LLMs, and use LLM as a Judge to check data against criteria that people define
- Once the goal and budget are agreed, the person in charge decides how to proceed and where to use AI
- Every role is expected to redesign its own work with AI. Our leadership team writes code too, and you discuss design decisions with them directly
### Development process
We run agile development based on Scrum. You build and improve production data pipelines in a mission-based Scrum team. We hold daily scrums, sprint planning, sprint reviews, and postmortems.
### Team
- With colleagues already working from Kyushu, you can build production data infrastructure while continuing to live where you are in Japan
- Almost-weekly discussions with leadership and business colleagues give you direct access to the customer problems behind data requirements, and you use that context to shape data models and quality criteria
## <Why Join>
### Build the data that people and AI agents work on
People use this data in SalesNow, or reach it through AI agents such as Claude and Cursor via SalesNow MCP. You design the entity matching and validation behind both. The platform holds 8 billion records gathered through three routes, including offline collection and user contributions that no crawler reaches. A matching rule that works on one route still has to be shown to hold on the others, and you decide what evidence settles that. You come away able to design entity resolution across sources of very different shapes.
### Define the requirements yourself
You decide what to collect, from which of our three routes, and under what conditions a record is fit to ship. When web crawling, offline collection, and user contributions give different answers for the same company, you set which evidence wins. For example, you define when a company URL found on the web can be linked to a Japanese corporate number. That rule decides what customers see in SalesNow and what AI agents receive through SalesNow MCP.
### Build the platform that carries the ARR-per-employee goal
We have set a goal to grow ARR per employee tenfold in three years from November 2025. On the data side, the platform you design is how we get there: by redesigning collection and verification, not by adding people to check records.
In two to three years, using AI will be a given. What will set people apart is the ability to redesign workflows and organizations around AI. Here, you sit close enough to the business to see which decisions the data has to support, and you have the room to change how the data team works with AI to support them.
## <Hiring Process>
1. Application screening
2. Interviews (multiple rounds)
3. Reference check
4. Offer meeting (offer)
## <Links>
- Recruiting site: https://recruit.salesnow.co.jp/
- Interview with a SalesNow engineer who joined as an intern and now develops SalesNow MCP and our API (Japanese): https://note.com/salesnow/n/na813feacc9cb
- Full job details (Japanese): https://herp.careers/v1/salesnow0801/aLtdW6WbCZ0h
必須スキル
This position is open to non-Japanese candidates who can work in Japan as full-time employees and use Japanese at a business level.## <About SalesNow>
SalesNow's mission is 「誰もが活躍できる仕組みをつくる。」 (Create systems where everyone can thrive), and we are taking on the challenge of fundamentally changing how people work. We rebuild how B2B work gets done, using corporate data and AI so that each person can create more value. SalesNow is the corporate data infrastructure for the AI era, supplying the data that people and AI agents work from. Our database covers over 16 million corporate and organizational profiles, structured down to individual offices and departments, and built on 8 billion records. SalesNow ranks No.1 in corporate database record count (企業データベース収録件数No.1; survey period: October 2025; research by Japan Marketing Research Organization).
SalesNow Data Network, our own collection network, combines three routes. Our platform uses AI to update web information daily. A partner network of tens of thousands of data researchers, working alongside our own in-house team, collects offline — from on-site surveys to reading paper documents. Our own corporate-information media draws 7 million page views a month, and that user network also takes on part of the offline collection. Together they give us what search engines never return, including org charts and more than 7.5 million department contacts.
Our own JC code assigns a unique ID to each site — headquarters, plants, branches, and stores — not only to the company. Multi-site groups can therefore be traced accurately. Records are matched on corporate number, name, address, phone, URL, and email. Fifteen levels of priority resolve them to a single company. For unresolved company matches, our team uses an AI workflow to supplement company information from current web sources. The platform can attach more than 50 attributes to company records, including registry information, industry classification, revenue, capital, headcount with its rate of change, sites, contacts, and scores. This collection, matching, and validation work is what you will design at production scale.
Sales teams use signals such as a rise in hiring, a funding round, a new office, or a stop in job postings to see why a company is worth approaching now. Customers use the data in the corporate database SalesNow, in their Salesforce and HubSpot records, and from AI agents such as Claude and Cursor through SalesNow MCP. Our Data API (beta) ships as four APIs: Company, News, Recruit, and Organization (departments, org charts, and sites). Companies that have adopted SalesNow include Yamato Transport, Panasonic, LY Corporation, SMBC Nikko Securities, PERSOL CAREER, JCB, and GMO Payment Gateway. You develop your data engineering expertise on a product with enterprise adoption, where the matching and validation you design decide what those customers can rely on.
## <Why We Are Hiring>
**The era when revenue grew in proportion to headcount is over.** As Japan's workforce shrinks, each person needs to create more value. We have set a goal to grow ARR per employee tenfold in three years, starting from November 2025. On the data side, that means redesigning how corporate data is collected, matched, and verified, rather than adding people to check it.
SalesNow is approaching Series B. You will set the collection requirements and the quality criteria behind SalesNow MCP and the Data API: which sources to trust, what evidence a record needs, and when a pipeline result is good enough to ship.
The same record now serves the SalesNow screen, customers' Salesforce and HubSpot data, the Data API, and AI agents through SalesNow MCP. We need an engineer who owns collection, matching, and verification end to end, because a quality rule set at one stage decides what all four of those consumers receive.
## <Mission>
You will build structured company profiles that capture what each company does and what has changed recently. You will tell apart companies with the same name and handle address variations, mergers, and relocations. You define the matching rules that distinguish separate companies and link records belonging to the same company. Customers rely on those records when they research companies in SalesNow, or through AI agents connected via SalesNow MCP.
## <What You Will Do>
You will build and improve our corporate data platform, which runs at the scale of tens of terabytes. You will take on one or more of the following:
- Build and improve pipelines for web data crawling
- Implement data cleansing and enrichment
- Build a monitoring environment to improve data quality
- Improve the data platform for scalability, performance, and cost efficiency
- Design the data platform and data quality foundations with future data science use in mind
### Examples of problems you will work on
- Decide whether a job posting is still open. Closure detection is currently split across three systems and behaves differently by source. Some sources do not show closures on their listing pages. If disappearance from a crawl triggers closure, a single missed page can incorrectly mark a posting as closed. You will set criteria for each source that balance false closure detections against missed closures
- Keep company website information current. You will design the process that fetches information from company websites and keeps tracking what has changed since the last fetch, so that each company's profile stays up to date
※Scope of changes to duties: yes (従事すべき業務の変更の範囲:有り)
## <Tech Stack>
- Data collection: Python (Scrapy), AWS Fargate for Amazon ECS, Amazon Managed Workflows for Apache Airflow
- Data storage: Amazon S3, Amazon Aurora PostgreSQL
- Data transformation: Scala and Python on Databricks (Apache Spark, Delta Lake; processed in three layers: raw, cleansed, and serving)
- Infrastructure: AWS, Databricks
- Other: GitHub, Slack, Notion, Asana, Google Workspace
- Development tools: Claude Code, Codex, Cursor, CodeRabbit, n8n
## <How We Work>
### Working with AI
- You use company-funded AI tools to build and verify pipelines; spending averages over JPY 100,000 per person per month
- Everyone can use Claude and Gemini. We also use Codex. You compare models on matching and extraction tasks, and select them based on your own quality checks
- You have several LLMs review your pipeline code and data rules, and you decide which findings to adopt after checking the evidence yourself
- Engineers use CodeRabbit for first-pass PR reviews, and focus their own review on design validity, data assumptions, impact, edge cases, data quality, and reproducibility
- In data processing, we structure listed companies' disclosure documents with LLMs, and use LLM as a Judge to check data against criteria that people define
- Once the goal and budget are agreed, the person in charge decides how to proceed and where to use AI
- Every role is expected to redesign its own work with AI. Our leadership team writes code too, and you discuss design decisions with them directly
### Development process
We run agile development based on Scrum. You build and improve production data pipelines in a mission-based Scrum team. We hold daily scrums, sprint planning, sprint reviews, and postmortems.
### Team
- With colleagues already working from Kyushu, you can build production data infrastructure while continuing to live where you are in Japan
- Almost-weekly discussions with leadership and business colleagues give you direct access to the customer problems behind data requirements, and you use that context to shape data models and quality criteria
## <Why Join>
### Build the data that people and AI agents work on
People use this data in SalesNow, or reach it through AI agents such as Claude and Cursor via SalesNow MCP. You design the entity matching and validation behind both. The platform holds 8 billion records gathered through three routes, including offline collection and user contributions that no crawler reaches. A matching rule that works on one route still has to be shown to hold on the others, and you decide what evidence settles that. You come away able to design entity resolution across sources of very different shapes.
### Define the requirements yourself
You decide what to collect, from which of our three routes, and under what conditions a record is fit to ship. When web crawling, offline collection, and user contributions give different answers for the same company, you set which evidence wins. For example, you define when a company URL found on the web can be linked to a Japanese corporate number. That rule decides what customers see in SalesNow and what AI agents receive through SalesNow MCP.
### Build the platform that carries the ARR-per-employee goal
We have set a goal to grow ARR per employee tenfold in three years from November 2025. On the data side, the platform you design is how we get there: by redesigning collection and verification, not by adding people to check records.
In two to three years, using AI will be a given. What will set people apart is the ability to redesign workflows and organizations around AI. Here, you sit close enough to the business to see which decisions the data has to support, and you have the room to change how the data team works with AI to support them.
## <Hiring Process>
1. Application screening
2. Interviews (multiple rounds)
3. Reference check
4. Offer meeting (offer)
## <Links>
- Recruiting site: https://recruit.salesnow.co.jp/
- Interview with a SalesNow engineer who joined as an intern and now develops SalesNow MCP and our API (Japanese): https://note.com/salesnow/n/na813feacc9cb
- Full job details (Japanese): https://herp.careers/v1/salesnow0801/aLtdW6WbCZ0h
歓迎スキル
- Data processing with Apache Spark (including PySpark)- Experience with AWS analytics services
- Experience with Databricks
- Experience in data management
- Alignment with DataOps principles
求める人物像
- You share our mission: 「誰もが活躍できる仕組みをつくる。」 (Create systems where everyone can thrive). You work by our values — AI Native, Kotoshikou (purpose-driven), and Highest Quality — and you intend to make full use of AI- You enjoy moving fast. Even as things change, you start from the purpose, find the problem, and see it through to a solution
- You work with your team and customers with integrity and ownership
- You take in new technologies and knowledge eagerly, and take on product development that drives business growth
- You take what customers and business teams tell you back into the data model and the quality criteria the product runs on
For details, see our Culture Deck: https://speakerdeck.com/salesnow/culture-deck
応募概要
給与
JPY 4,000,000 – 15,000,000 per year勤務地
Within Japan (fully remote). We have an office in Shibuya.※Scope of changes to place of work: none (就業場所の変更の範囲:無し)
雇用形態
Full-time勤務体系
### Work style- Fully remote within Japan
- Near Tokyo: one recommended office day per month
- Outside the Tokyo area: attend company-wide events once or twice a year only
### Working hours
- Full flextime (no core hours)
- Standard working hours: 8 hours per day
### Holidays and leave
- Two days off every week (Saturdays, Sundays, and national holidays)
- Year-end and New Year holidays
- Summer holidays
- Paid leave (10 days granted after 6 months; up to 20 days can be carried over)
- Female leave (5 days per year)
- Special leave for weddings and funerals
- Maternity leave and childcare leave
- Family care leave
試用期間
3 months福利厚生
### Allowances- Commuting costs fully covered
- Book allowance (up to JPY 5,000 per month)
- Skill development and productivity support (up to JPY 10,000 per month)
- Socializing allowance (飲みニケーション手当)
- Referral hiring allowance
### Health and support
- Social insurance (health insurance, employees' pension, workers' accident compensation insurance, employment insurance)
- Babysitter subsidy
- Comprehensive medical checkup (every other year, age 40 and over)
### Other
- Stock option plan