Vai al contenuto principale
JobCannon
Tutte le competenze

Apache Iceberg Data

⬢ LIVELLO 3Tecniche
Alto
Impatto sullo stipendio
6 mesi
Tempo di apprendimento
Medio
Difficoltà
—
Carriere
In sintesi

Apache Iceberg is Netflix's open-source table format that brings ACID transactions, time-travel, and schema evolution to data lakes. Unlike Hudi (format-agnostic but opinionated), Iceberg is format-focused and works across Spark, Flink, Presto, Trino, and BigQuery. Advanced practitioners leverage Iceberg's hidden partitioning, partition pruning, and evolution for analytics workloads. Iceberg is rapidly becoming the standard for open data platforms. Salary impact: $120k-180k for senior Iceberg engineers in data-heavy companies.

Cos'è Apache Iceberg Data

Apache Iceberg is an open table format that brings ACID transactions, time-travel, and schema evolution to data lakes. Created by Netflix, Iceberg separates the physical file layout from the logical table definition. This enables: - ACID semantics: Consistent reads and writes across partitions

🔧 STRUMENTI ED ECOSISTEMA
IcebergSparkFlinkTrinoPrestoDuckDBS3HDFSGlue

💰 Stipendio per regione

RegioneLivello baseMidLivello esperto
USA$95k$140k$210k
UK£65k£105k£155k
EU€70k€110k€165k
CANADAC$105kC$150kC$230k

❓ Domande frequenti

How does Iceberg compare to Delta Lake?
Delta is Databricks-centric, supports Spark natively. Iceberg is format-neutral, works with Spark, Flink, Presto, BigQuery. Iceberg's hidden partitioning and partition evolution are superior. Delta has stronger Databricks ecosystem. For open platforms, Iceberg is the better choice.
What are hidden partitions in Iceberg?
Users query without knowing partition scheme. Iceberg handles partitioning transparently via partition transforms (year, month, day, bucket). Enables flexible partitioning without affecting queries.
How does Iceberg handle schema evolution?
Columns can be added, removed, renamed, or reordered. Iceberg tracks schema changes in versioned metadata. Reads respect schema evolution, old columns null-filled if removed. No backfilling needed.
Can I query Iceberg tables with BigQuery?
Yes, Iceberg is a BigQuery-supported open format. BigQuery can directly query Iceberg tables on GCS. No ETL needed.
Is Iceberg suitable for streaming?
Yes, via Flink and Kafka. Iceberg + Flink enables exactly-once streaming. Slower write throughput than raw Parquet but guarantees consistency.
How does Iceberg handle partitioning evolution?
Change partitioning scheme without rewriting data. Old data stays in old partition layout, new data uses new scheme. Metadata tracks both. Gradual evolution possible.
What's the query performance difference vs. raw Parquet?
Iceberg metadata overhead is ~5-10% for OLAP queries. Metadata pruning can be faster (skips files without reading). For OLTP (point lookups), no advantage. Metadata cost is negligible for large table scans.

Non sei sicuro che questa competenza faccia per te?

Fai il Career Match — ti suggeriremo i percorsi giusti.

Trova le competenze adatte a te →

Trova il tuo percorso di carriera ideale

Abbinamento basato sulle competenze per 2521 carriere. Gratis, ~3 minuti.

Fai il Career Match — gratis →