我在 2026 年 3 月走了 JPMC data engineer 的面试 loop,职位是在 NYC 的 commercial banking data platform 团队。分享细节,因为我去之前几乎找不到任何有用的信息。
整个流程一共五轮。先是 recruiter 的 phone screen(comp expectations、时间线、常规问题)。然后是和团队里的 senior DE 的 technical screen。接着是三轮 back to back 的 onsite,分两天,他们是 virtual 做的。
Technical screen 60 分钟。一半 SQL,一半“walk me through a pipeline you built and what broke.(带我过一遍你做过的一个 pipeline,以及它哪里出过问题。)”。SQL 真的很难。window functions、CTEs,还有一道涉及 ranking with ties 的题。不是 LeetCode 中等难度,这是实打实的 query optimization 思维。我得解释为什么选某种 join order,以及 query planner 会怎么做。
Onsite round 1: system design for data. 给了个场景:设计一个实时交易数据的 ingestion layer,用来喂给 fraud detection model。我讲了 Kafka、partitioning 策略、schema registry、怎么处理 late-arriving data。他们对我的 retry logic 反驳得很凶。准备好为你的选择辩护。
Onsite round 2: more SQL and Python. SQL 还是 window functions。Python 是 pandas,然后他们又切到“how would you do this in PySpark(你会怎么用 PySpark 来做这个。)”来做其中一个 transformation。我现在的工作里 Spark 用得不多,这点有点露怯。
Onsite round 3: behavioral. 标准 JPMC behavioral 格式。三个 STAR-method 问题。“Tell me about a time you caught a data quality issue in production.(讲讲一次你在生产环境中发现数据质量问题的经历。)” “Describe a time you had to influence a stakeholder without authority.(描述一次你在没有正式权限的情况下影响 stakeholder 的经历。)” 经典题,但他们想要具体细节。
有几件事比我预期更关键:理解 data lineage;真正懂 batch 和 streaming 架构的差别,不只是表面;以及能讲清楚你在 pipelines 上怎么做 observability(不只是能不能跑,还要怎么知道数据是对的)。
从申请到 offer 的总时间线大概 6 周。offer 在 NYC 的 senior DE band 里算合理,不过我也要说 JPMC 并不会去 match 纯 tech 公司的 comp。如果你想知道具体数字,去这家公司页面的 comp thread 看。