Ten years may not sound like much, but in the world of tech, it’s a lifetime. While updating my CV, I realized that 8 years have already passed since my first introduction to Databricks, and a full decade since I set foot in the industry. It felt like the perfect moment to look back on how far we’ve come and how this field has transformed before our eyes.
The yellow elephant in the jungle
My career began in the Hadoop era and I witnessed the shift from on-premise solutions to the cloud.
In those days, many tools were emerging, and it wasn’t easy to keep up with their names and functionalities. What made it even more complicated was the widespread reliance on non-native cloud components. For example, for data ingestion, solutions running on a virtual machine were often still preferred over solutions like Azure Data Factory, since it wasn’t yet the powerful data integration service it is today. For instance, I regularly used tools like Apache Oozie and Sqoop, both of which have now found their way into the Apache Attic.

In 2017, I found myself attending one of the early Databricks-on-Azure sessions for Microsoft partners. At the time, it felt like just another player in a crowded Big Data scene. I recall that one of the key takeaways was a simple yet groundbreaking slide that finally made sense of Azure’s data services. For the first time, all the disparate Azure data products were neatly grouped on a single slide, categorized by their primary use.
Honestly, I went home from the session not particularly excited. While Databricks clearly offered a very powerful solution, similar to other Spark-based platforms, with impressive notebook collaboration features, it felt a bit dry in other capabilities.
Looking back at that period, when “data was the new oil”, only one solution seemed to be the undisputed champion and silver bullet: Apache Hadoop.
Hadoop, however, often felt cumbersome. Its various services seemed to fight rather than work together, and attempts to create managed services were never truly convincing. For example, Azure HDInsight was a good attempt at a managed Hadoop cluster, but it wasn’t 100% reliable. It often required multiple restarts and could be overly complex to manage. Databricks, on the other hand, offered far superior cluster stability but initially lacked robust orchestration capabilities, especially given the limited integration with Azure Data Factory v1.
Before Databricks reached General Availability (GA), I changed projects and ended up with a hybrid Data Platform implementation: HDInsight in the cloud and Hortonworks on-premise. For a while, Databricks went out of my sight.
The red bricks are back
By 2020, Databricks re-entered my world — and this time, it had a glow-up. Delta Lake was technically “there,” but not yet the centerpiece. What did catch my attention was how polished everything had become and the troubleshooting features: Ganglia provided all the information needed to check what was happening.
Then came the real shift: the Lakehouse.
Delta Lake existed earlier but it wasn’t until the launch of Delta Lake 1.0 in 2021 that things truly took off. Adoption surged. Buzz turned into real-world impact. The architecture clicked and the narrative itself proved to be a game-changer.
The cherry on top was the introduction of Unity Catalog, which finally enabled centralized data governance, reminiscent of the old-school Apache Ranger. What had once been “just another Spark platform” had quietly become the future.
Fast forward to today, Databricks has truly evolved into a comprehensive data platform, a data platform on steroids, now equipped with robust orchestration and even unexpected dashboarding features.
With every Data + AI Summit, the feature list grows so fast it’s hard to keep up. But what really has my attention right now is Lakebase, a true serving layer directly within the platform.
The new urban jungle

Looking back at a decade in data engineering, my Databricks journey feels like a microcosm of the entire industry’s evolution. What began as a powerful Spark-based solution became a leading force in the analytics space, simplifying the work of data engineers and other data practitioners.
Gone are the days of wrangling Linux boxes, juggling fragile toolchains, and duct-taping workflows together. Delta Lake was the tipping point, but the real magic is in the platform’s ability to abstract away the chaos, which frees data engineers to focus on value.
Databricks’ relentless pace of innovation proves one thing: the learning curve may remain steep, but the impact curve is exponential.
As data professionals, our constant adaptation remains the key. Where we once constantly chased new technologies emerging across the landscape, it’s now about going deep within ever-expanding platforms.
We’ve moved past the age of the nitty-gritty of infrastructure. The real value now lies in how we generate insight, build trust, and ship outcomes at speed.
But here’s the twist: as platforms simplify, automate, and abstract — will data architecture lose its edge, or will it evolve into something even more critical? Time will tell.
But the challenge for you is now: are you riding the wave or watching it pass you by?
