Hoppa till huvudinnehåll
JobCannon
Alla kompetenser

Apache Iceberg Data

⬢ NIVÅ 3Tekniskt
Hög
Lönepåverkan
6 månader
Tid att lära sig
Medel
Svårighetsgrad
—
Karriärer
I korthet

Apache Iceberg is Netflix's open-source table format that brings ACID transactions, time-travel, and schema evolution to data lakes. Unlike Hudi (format-agnostic but opinionated), Iceberg is format-focused and works across Spark, Flink, Presto, Trino, and BigQuery. Advanced practitioners leverage Iceberg's hidden partitioning, partition pruning, and evolution for analytics workloads. Iceberg is rapidly becoming the standard for open data platforms. Salary impact: $120k-180k for senior Iceberg engineers in data-heavy companies.

Vad är Apache Iceberg Data

Apache Iceberg is an open table format that brings ACID transactions, time-travel, and schema evolution to data lakes. Created by Netflix, Iceberg separates the physical file layout from the logical table definition. This enables: - ACID semantics: Consistent reads and writes across partitions

🔧 VERKTYG & EKOSYSTEM
IcebergSparkFlinkTrinoPrestoDuckDBS3HDFSGlue

💰 Lön per region

OmrådeNybörjareMidErfaren
USA$95k$140k$210k
UK£65k£105k£155k
EU€70k€110k€165k
CANADAC$105kC$150kC$230k

❓ Vanliga frågor

How does Iceberg compare to Delta Lake?
Delta is Databricks-centric, supports Spark natively. Iceberg is format-neutral, works with Spark, Flink, Presto, BigQuery. Iceberg's hidden partitioning and partition evolution are superior. Delta has stronger Databricks ecosystem. For open platforms, Iceberg is the better choice.
What are hidden partitions in Iceberg?
Users query without knowing partition scheme. Iceberg handles partitioning transparently via partition transforms (year, month, day, bucket). Enables flexible partitioning without affecting queries.
How does Iceberg handle schema evolution?
Columns can be added, removed, renamed, or reordered. Iceberg tracks schema changes in versioned metadata. Reads respect schema evolution, old columns null-filled if removed. No backfilling needed.
Can I query Iceberg tables with BigQuery?
Yes, Iceberg is a BigQuery-supported open format. BigQuery can directly query Iceberg tables on GCS. No ETL needed.
Is Iceberg suitable for streaming?
Yes, via Flink and Kafka. Iceberg + Flink enables exactly-once streaming. Slower write throughput than raw Parquet but guarantees consistency.
How does Iceberg handle partitioning evolution?
Change partitioning scheme without rewriting data. Old data stays in old partition layout, new data uses new scheme. Metadata tracks both. Gradual evolution possible.
What's the query performance difference vs. raw Parquet?
Iceberg metadata overhead is ~5-10% for OLAP queries. Metadata pruning can be faster (skips files without reading). For OLTP (point lookups), no advantage. Metadata cost is negligible for large table scans.

Osäker på om den här kompetensen passar dig?

Gör Career Match — vi föreslår rätt spår för dig.

Hitta mina bäst passande kompetenser →

Hitta din ideala karriärväg

Kompetensbaserad matchning mot 2 521 karriärer. Gratis, ~3 minuter.

Gör Karriärmatchningen — gratis →