Palo Alto Networks · Primly 社区

Palo Alto Networks data engineer 面试:pipeline、SQL,以及一道真的很难的分布式系统题

analyst_ana (Primly starter) · 4 条回复

Did the PANW DE loop for a role on their cloud telemetry platform team. This was a senior DE position in Santa Clara, late 2025, so 2026 readers should expect similar structure.

Five rounds total: recruiter, coding screen, then a 4-round onsite.

Phone coding screen Leetcode-style SQL: given a table of security alerts, write a query to identify devices that triggered more than 3 distinct alert types within a 1-hour rolling window. It's a window function + GROUP BY problem. Took me 25 minutes, which they said was about average.

Onsite: Data pipeline design Design a streaming ingestion pipeline for firewall log data. ~100GB/day across thousands of enterprise endpoints. They asked about: Kafka vs Kinesis trade-offs, schema evolution (how do you handle a new log field without breaking downstream jobs), exactly-once semantics, and how you'd handle backpressure. This is where the security context matters. They process petabytes of telemetry and they want someone who's actually designed at that scale.

Onsite: SQL + transformation logic Two questions. One aggregation (median per group without a MEDIAN function), one deduplication problem (find the canonical record when you have duplicate device check-ins within a 30-second window). Straightforward if you've done data quality work before.

Onsite: Distributed systems This one surprised me. They asked: you have a real-time threat detection job reading from Kafka. The job crashes. When it restarts, how do you ensure you don't reprocess events that already triggered alerts AND don't miss any? We got into consumer group offsets, idempotent writes, and exactly-once processing. It's not a typical DE interview question; it felt closer to SRE territory. I think they care because a duplicate alert can create real operational noise.

Onsite: Behavioral Classic: tell me about a pipeline you owned that went down in production. How you detected it, how you communicated, what you changed. They want blameless post-mortem thinking.

Offer came in around $195k total, base $148k. 4-year vest. Pipeline took 6 weeks.

4 条回复

infra_ines (Primly starter)

那个 exactly-once Kafka 题也太硬了。他们是想让你写代码,还是只是偏设计、口头讨论?

由 AI 翻译,查看原文

de_derek (Primly starter)

纯口头 design,但他们对每一层都 push back。我说到“transactional writes”,他们就追问我具体指的 transaction boundary 是什么。每一层都要准备好能讲清楚、能 defend。

由 AI 翻译,查看原文

backend_bekah (Primly starter)

schema evolution 这个问题我面试里经常答崩,因为我会聊 Avro,然后默认他们知道我在说什么。他们是想让你讲到具体细节,还是保持概念层面就可以?

由 AI 翻译,查看原文

sec_sasha (Primly starter)

telemetry team 关注重复告警抑制太合理了。吵闹的 SIEM 是安全运维里最大的头疼之一。他们面试会考这个,是个很好的信号。

由 AI 翻译,查看原文