Warn is the default and filters nothing: how to go from scattered checks to a quality system in Databricks, with rules loaded dynamically from a Delta table, a quarantine…
How to query Postgres and MySQL from Databricks without moving data: connections and foreign catalogs in Unity Catalog, which part of the query gets pushed down to the…
What Liquid Clustering is, why it replaces partitioning and Z-ORDER, the CLUSTER BY syntax (including AUTO), how it gets triggered by OPTIMIZE, and a reproducible lab where…
What OpenSharing inherits from Delta Sharing and what it adds: true zero-copy, shares to Iceberg clients like Snowflake and Trino, models and agent skills as shareable…
What Databricks’ vectorized engine is, where it runs, what it speeds up and what it doesn’t, how to measure how much Photon your query uses, a with/without Photon benchmark…
How to deploy a complete RAG agent with memory, guardrails and CI/CD using a single databricks.yml — with real code, step by step.
Detect, Assess, Remediate, Verify — how the Databricks agent that monitors, diagnoses, and proposes fixes for your production pipelines works.
A curated guide to the most useful repositories in the Databricks GitHub ecosystem: from the four official orgs to community tools and MVPs. With context on what each thing…
LTAP unifies OLTP and OLAP on a single copy of data over Delta and Iceberg. What it is, how it works, how it differs from HTAP and Zero ETL, and a concrete example of how it…
Genie Code specializes in ML engineering, Lakeflow Designer leaves preview, AI Runtime adds multi-node training, and Nadella records a fireside chat with Ghodsi. Recap of…
Lakehouse//RT, LTAP, Genie One, CustomerLake, Lakewatch, OpenSharing and more. A complete recap of Data + AI Summit 2026 from afar, with links to every announcement.
Compose, govern and share: Omnigent is the missing layer on top of Claude Code, Codex and Pi. Installation, YAML configuration, policies and a real test.
Streaming tables, materialized views, expectations, AUTO CDC with SCD Type 1/2, triggered vs continuous modes, serverless, and the gotchas that aren’t in the tutorial.
Rate limiting, guardrails, usage tracking, cost attribution, fallbacks and traffic splitting. Everything you need to govern LLMs across your organization.
Classic, Pro or Serverless: which one to pick, how to size it, and the gotchas that will save you money. With Photon, Query Federation, AI Functions and real costs.
File arrival triggers, table update triggers, Trigger.AvailableNow, Job Clusters and the limitations that aren’t in the tutorial.
Databricks Container Services, custom images, golden containers, CI/CD with Docker, and the mistakes that will cost you hours.
Data Mesh beyond the hype. The 4 principles, a real implementation with Unity Catalog, and why most companies get it wrong.
What Data Contracts are, why you need them, and how I implemented a framework in production with Databricks and PySpark.
A practical comparison of the three most widely used data models in the industry. With real examples, trade-offs, and a guide for choosing.
Feature Store, point-in-time lookups, online features, feature freshness, and the classic data leakage mistakes.
Model registration in UC, aliases instead of stages, data-to-model lineage, and model serving with AI Gateway.
Trigger modes, watermarks, foreachBatch, Auto Loader, and the traps that make you lose data or money.
Three-level namespace, inherited GRANTS, automatic lineage, row/column security and common governance mistakes.
Liquid clustering, OPTIMIZE, Z-ORDER, vacuum, time travel and tricks that change how you work with Delta.
Advanced use cases, deployment patterns, complex variables and tricks I picked up working with DABs in production.
Python, SQL, Spark, dbt, Databricks — the order matters. What I learned after years of working with data and what I’d do differently if I had to start over.
The 8 ingestion patterns from Bartosz Konieczny’s book. Examples in PySpark/SQL and how they map to the Medallion architecture.
What DataOps really is, how it differs from DevOps, the three pillars (automation, observability, collaboration), and common mistakes in production.
Review of the book by Joe Reis and Matt Housley (O’Reilly, 2nd ed.). The data lifecycle, key architectures, undercurrents and what changed in this edition.
Why we need a Data Engineering podcast in Spanish. Databricks, data architecture, MLOps and career — from the region, for the region.