上个月刚走完 VMware data engineer 的面试 loop,senior 级别,remote,cloud infrastructure data team。发出来是因为网上几乎找不到他们 DE 流程的有用信息。
recruiter screen 之后是四轮:
Round 1:SQL + data modeling,45 分钟。 这轮才是真正的筛子。他们先给了一个自己设计的 schema(有点像简化版 vCenter telemetry:hosts、VMs、resource pools),让我写 window functions 算 running averages、识别 outliers,还要做多表 join,并且有一些很刁钻的 NULL 处理。不是 Leetcode 那种,更像「这是个真实 schema,证明你会用」。他们还明确说用这种方式考 SQL,因为更贴近工作。
然后让我为一个新的 ingestion use case 设计数据模型。我画了 medallion architecture(bronze/silver/gold),并讲了 partitioning 策略。他们对我的 partitioning 选择稍微 push back 了一下,感觉更像技术讨论,而不是故意刁难。
Round 2:Python + pipelines。 在共享 IDE 里写代码。有个任务是写一个 pipeline:从 Kafka topic(mock 的)读数据,transform records,然后写到目标端并去重。他们很在意 error handling 和 idempotency。我一直强调「如果这一步执行到一半失败会怎样」,面试官反应挺好。
Round 3:面向数据的 system design。「Design a pipeline to ingest VM performance metrics at scale。」经典题。聊了 Kafka、Flink/Spark Streaming、Delta Lake、monitoring。他们问了很多 late-arriving data 和 exactly-once semantics。面试官很有主见,感觉像来回讨论,不像答题。
Round 4:Behavioral。 直接 STAR。跨团队冲突、我 push back 一个技术决策的经历、一个我会重新做的项目。
大概十天后给了 offer。流程很有条理,反馈也是真反馈。最需要重点准备的是 SQL 那轮。window functions、CTEs,还有准备好聊你的 indexing 选择。