Databrick Data Engineer

Hace 2 meses

República Dominicana Lumenalta Jornada completa

At Lumenalta, we partner with forward-thinking organizations to build technology solutions that scale, delight users, and accelerate business growth. Our global teams bring curiosity, commitment, and technical excellence to every project. We value transparency, autonomy, and impact—empowering every team member to do their best work.


We're seeking an experienced Data Engineer / Data Analyst to join a high-performance motorsport client building a real-time Databricks Lakehouse platform. As the client moves high-frequency telemetry, optical tracking, and timing/scoring feeds into the Lakehouse, they need someone who can establish the foundation of trust the entire platform depends on: deeply profiling raw data as it lands, and defining the data quality rules that decide when data is clean, complete, and reliable enough to be promoted downstream. Your work is the difference between data that engineers, analysts, and race officials can act on with confidence—and data they can't. You'll partner closely with those teams to define what "trustworthy" means for each feed, then encode those definitions directly into the pipeline.


Actively Hiring

We are hiring for a current opening on an active client project. This is a specific, presently open role. We review applications on a rolling basis and aim to move qualified candidates through our process promptly.


What You’ll Be Doing

  • Profile high-frequency, multi-sensor data as it lands in the Bronze layer—characterizing distributions, null and completeness rates, value ranges, cardinality, uniqueness, referential integrity, schema drift, duplicate and late/out-of-order records, and sensor dropouts across telemetry, optical tracking, and timing/scoring feeds.
  • Partner with business and domain stakeholders—race engineers, analysts, and competition officials—to understand how each feed is used and translate their domain knowledge into concrete, testable expectations: what "valid," "complete," and "in range" actually mean for telemetry, timing, and optical tracking data.
  • Define, document, and codify the data quality rules and expectations—validity, completeness, uniqueness, timeliness, range and threshold checks—that govern the promotion of data from Bronze to Silver.
  • Build and maintain the Bronze-to-Silver transformation logic: schema enforcement, type casting, deduplication, standardization and conforming, and quarantine of records that fail quality checks—producing clean, trusted Silver tables.
  • Implement data quality enforcement using Databricks-native tooling (Delta Live Tables / Lakeflow Declarative Pipelines expectations, or frameworks such as Great Expectations or Databricks DQX), including quarantine tables and alerting for records that don't meet the rules.
  • Establish data contracts and clear documentation of expected schemas and semantics for each source, working with upstream data producers and downstream consumers to keep expectations aligned as feeds evolve.
  • Monitor data quality over time—tracking rule pass/fail rates, surfacing data drift and anomalies within the pipeline, and alerting the team when source data or ingestion quality degrades.
  • Work within the Unity Catalog governance framework to ensure Silver datasets are trusted, lineage-tracked, and analysis- and model-ready for downstream engineering and data science functions.
  • Collaborate with data engineers to define the models, conformed schemas, and serving-layer outputs that clean Silver data must support—so downstream analytics and predictive race analysis are built on reliable inputs.


What We’re Looking For

  • 4+ years as a Data Engineer, Analytics Engineer, or data-focused Analyst, with a track record of delivering production-grade data pipelines and data quality workflows in data-rich environments.
  • Databricks certifications are required (Databricks Certified Data Engineer Associate or Professional preferred).
  • Strong data profiling skills—able to characterize unfamiliar, high-volume datasets using descriptive statistics, distribution analysis, and null, cardinality, uniqueness, and integrity checks.
  • Hands-on experience defining and implementing data quality rules and expectations, with familiarity in frameworks such as Delta Live Tables / Lakeflow expectations, Great Expectations, or Databricks DQX.
  • Solid understanding of the medallion architecture (Bronze / Silver / Gold) and Bronze-to-Silver promotion patterns: sch