今年早些时候走了 Goldman Sachs data engineering 的面试 loop,Engineering Division,NYC 的职位。recruiter screen 之后总共四轮。这里是拆解。
recruiter screen:30 分钟,主要是过简历,然后问了几个 data modeling 的高层问题。他们问我是否做过「large-scale data pipelines」,意思是要你说具体数字:每天多少行、延迟 SLA、一些能落地的指标。
round 1,coding + SQL:一道中等复杂度的 coding(Python,图遍历,不算变态)和两道 SQL。SQL 重点在 event log 的 partitioning 和 aggregation。有一道故意把需求写得含糊,他们想看你在写之前会不会先澄清。
round 2,面向数据的 system design:设计一个实时 trading data pipeline。从 market feeds 摄取,为低延迟查询和历史分析分别存储。他们追得很凶:feed 掉了怎么办,怎么保证 exactly-once delivery,backfill 怎么做。kafka 立刻就提到了。他们还追问 consumer group lag monitoring,我刚好真的做过,所以比较顺。
round 3,behavioral:这个让我意外,是完整的 45 分钟 behavioral。STAR format。很多「tell me about a time you disagreed with a technical decision(讲一次你不同意某个技术决策的经历)」和「how did you handle a production incident.(你是怎么处理一次生产事故的)」。GS 对这个的权重比我以为的高,尤其对 eng 角色。
round 4,technical deep dive:一个 senior IC 把我过去的一个项目挖得非常细。我要解释我两年前做过的架构选择。他们是真的好奇,不是想坑我。
pipeline + systems design 是决定生死的一轮。如果你至少在理论上没有端到端设计过 streaming data system,去面之前先把这个补上。