Data Engineering is a small, high-leverage platform team serving every Sunday entity (insurance, broker, care, technology) across two countries. We run one lakehouse and a set of frameworks — not a pile of one-off pipelines. You will have unusual ownership: the systems you design run production for the whole group.
Key Areas of Responsibility:
- The lakehouse: Databricks on AWS (Unity Catalog, Delta) across Thailand and Indonesia, prod and non-prod workspaces.
- Ingestion & export frameworks: our Python/PySpark (Spark Connect) batch framework pulling from operational databases (Postgres, MySQL, MSSQL, DB2, MongoDB) into Delta — and pushing back out: hash-diff incremental reverse-ETL from the lakehouse into operational Postgres serving live systems, plus Kafka (Avro/Confluent) streams.
- dbt warehouses: layered medallion architecture for Thailand and Indonesia, tested and CI-gated.
- Data Governance : End to End data governance across Thailand and Indonesia.
- Everything as code: Terraform/Terragrunt for infrastructure, Databricks Asset Bundles for jobs and schedules, Jenkins for CI/CD. If it isn't in git, it doesn't exist.
Responsibilities:
- Design, build, and operate batch and incremental pipelines end-to-end — from source system to the table an underwriter's dashboard reads.
- Extend frameworks rather than write one-offs: when you solve a problem, the next ten sources get the solution for free.
- Own data quality and reliability: tests, monitoring, SLAs, and the root-cause analysis when something breaks at 7am.
- Model insurance data (policies, claims, members, payments) so analysts and data scientists can trust and reuse it.
- Partner directly with data scientists, analysts, and business stakeholders across four business entities and two countries.
- Review more code than you write — including code written by AI. We work AI-assisted by default.
- Mentor other engineers and raise the team's bar for design and operational discipline.
Requirement:
- 5+ years building production data systems, with at least one system you owned end-to-end (design → build → operate → evolve).
- Expert SQL and solid data modeling — dimensional and lakehouse/medallion patterns, and the judgement to know when each applies.
- Strong Python engineering: typed, tested, reviewable code — not just notebooks.
- Production experience with Spark and a lakehouse platform (Databricks strongly preferred) or equivalent scale elsewhere.
- Real understanding of incremental processing: CDC, merge/upsert semantics, idempotency, late-arriving data, backfills.
- Operational depth: you've debugged pipelines under pressure and can tell the story of a root cause you found.
- Fluency with AI coding tools and a track record of catching their mistakes. We don't screen AI out of our hiring process — we screen for people who use it well.
- Clear written and spoken English — it's our working language, and much of our design work happens in documents and PRs.