Genentech · Primly 社区

Genentech data engineer 面试:pipeline和SQL(走了senior DE岗位的全流程)

de_derek (Primly starter) · 7 条回复

几个月前走了Genentech的DE loop。我之前看到的内容大多很泛或者偏SWE,所以我写一下data engineer版本。

岗位: Senior Data Engineer,Research Informatics。South San Francisco,hybrid。

流程: Recruiter call -> HM call -> Technical screen(1小时)-> Full loop(4小时,现场;远程候选人则是视频)

Technical screen: 这是最偏实战的一轮。一道SQL题 + 一段pipeline design讨论。SQL不简单:多表join、window functions、处理临床数据集里的缺失值。他们特别问了在patient data语境下怎么处理NULL,这和常规分析里的NULL关注点不太一样。要想清楚:临床测量缺失是没测、未知,还是确实不存在。

Full loop: SQL和数据建模(60分钟):两题。一题schema设计,一题query优化。他们给了我一个很乱的schema,问我怎么为下游BI重构。 Pipeline / 架构(45分钟):让我设计一个基因测序数据的ELT pipeline。数据量很大(每批次按TB算),转换逻辑复杂,而且数据受监管。我讲了Airflow做编排、dbt做转换、Snowflake或BigQuery做数仓,并且花了很多时间在数据质量检查和日志记录上。他们对这部分感觉很在意。 Behavioral(45分钟):标准STAR题,但带一点领域变化,比如“tell me about a time your pipeline produced incorrect results and how you caught and fixed it.(讲一次你的pipeline产出了错误结果,你是怎么发现并修复的。)”。

他们看重的工具: Python、SQL、Airflow、dbt,至少熟悉云存储(S3/GCS)。他们对Snowflake和BigQuery没有明显偏好,但会让我对比。

薪酬(我的offer,因为地点原因拒了): base大概155k,15% bonus target,RSU大概45k/4yr。还行,但比同级别的纯tech公司低。

由 AI 翻译,查看原文

7 条回复

analyst_ana (Primly starter)

临床数据里对 NULL 的处理细节,这个点太好了,我之前完全想不到。在标准分析里 NULL 通常就是缺失/未知,但在临床语境下,「没采集」「不适用」「真正未知」的区别是有实际含义的。

由 AI 翻译,查看原文

ds_dmitri (Primly starter)

他们在 dbt 这块会问多深?比如会问 models、tests、sources,还是更偏表面?

由 AI 翻译,查看原文

de_derek (Primly starter)

深度中等。他们问了我会怎么组织模型(staging vs marts),以及我会不会用 dbt tests、会用哪些。我提了 schema tests 和针对临床数据检查的自定义 generic tests。他们看起来挺满意的。没深入到 macros 或更高级的 Jinja 细节。

由 AI 翻译,查看原文

finance_faye (Primly starter)

你分享的 comp 跟我朋友去年类似岗位拿到的差不多。RSU 包比纯 tech 薄一点,但从稳定性角度看 Roche 这个母公司挺靠谱。算账方式不一样。

由 AI 翻译,查看原文

infra_ines (Primly starter)

Airflow 在大规模场景下真是又爱又恨。他们有问 Airflow 里关于失败处理和重试的具体问题吗,还是更通用的编排概念?

由 AI 翻译,查看原文

de_derek (Primly starter)

主要是一些通用概念:idempotency、backfill 策略、失败告警。我提了 SLA、DAG timeout,这些反馈不错。他们不算在跑最前沿的基础设施,所以我觉得他们不是在找很深的 Airflow internals。

由 AI 翻译,查看原文

Primly Team

One stage people tend to underestimate in senior DE loops is the assumption-setting before you write any SQL or draw an architecture box. In domains with high-stakes data (in most processes at this level), interviewers are often testing whether you pause to define semantics and guardrails, not just whether you can query.

A practical structure that works well: Restate the business question and the grain (patient, visit, sample, run). Enumerate “unknown vs not collected vs not applicable” and how you will encode each (NULL, sentinel, separate status column). Call out data quality checks you would enforce (uniqueness, referential integrity, outlier rules) and where they live (ingest vs transform vs serving). Only then write SQL or design the pipeline, and explain tradeoffs (performance vs traceability).

What’s one assumption you surfaced in a DE interview that changed your solution, especially around missingness or entity grain?