All Jobs Vacancy

Data Architect- Databricks

Posted 1 day ago by cloudingest inc

Key Responsibilities

Databricks Engineering

Build and optimize ETL/ELT pipelines using PySpark, Spark SQL, Databricks Workflows, and Delta Lake.

Develop Bronze/Silver/Gold Medallion architecture pipelines.

Implement Delta Live Tables (DLT) for automated ingestion and transformation.

Manage and optimize Databricks clusters, jobs, notebooks, repos, and workflows.

Perform Spark performance tuning (partitioning, caching, AQE, broadcast joins).

Implement Unity Catalog governance (catalogs, schemas, tables, permissions).

Integrate Databricks with AWS S3 / Azure Data Lake / Kafka / APIs.

Snowflake Engineering

Design and develop Snowflake ELT pipelines using Snowflake SQL, Streams, Tasks, and Snowpipe.

Build warehouse models, fact/dimension tables, and curated datasets.

Optimize Snowflake performance (clustering, micro-partitioning, query tuning).

Implement RBAC, masking policies, row-level security, and governance.

Integrate Snowflake with Fivetran, DBT, ADF, Glue, Kafka, or custom ingestion frameworks.

Data Pipeline & Integration

Build scalable ingestion frameworks for structured, semistructured, and unstructured data.

Integrate data from databases, APIs, cloud storage, streaming sources, and enterprise systems.

Implement data quality checks, validation rules, reconciliation, and lineage.

Support ML/AI workloads, feature engineering, and model-ready datasets.

Cloud & DevOps

Work with AWS, Azure, or Google Cloud Platform cloud-native services.

Implement CI/CD using GitHub Actions, Azure DevOps, GitLab, Jenkins.

Containerize workloads using Docker and orchestrate with Kubernetes (nice to have).

Monitor pipelines using CloudWatch, Azure Monitor, Databricks metrics, PrometheGrafana.

Required Skills

Core Technical Skills

Databricks (PySpark, Spark SQL, Delta Lake, Workflows, DLT)

Snowflake (SQL, Streams, Tasks, Snowpipe, RBAC)

Python

Advanced SQL

Cloud platforms: AWS / Azure / Google Cloud Platform

Data modeling (Star schema, dimensional modeling)

ETL/ELT pipeline development

Data quality, governance, lineage

CI/CD pipelines

API integration & REST services

Nice-to-Have Skills

DBT

Kafka / Spark Structured Streaming

MLflow / Feature Store

Terraform

Airflow / ADF / Glue / Dataflow

RAG/LLM data preparation (bonus)

Healthcare, finance, or regulated industry experience

Rate:
Not specified
Location:
Remote
IR35 Status:
Not specified
Remote Status:
Remote
Industry:
Data & Analytics
Seniority Level:
Not Specified

Take-Home Pay

Not Available

Visit calculators for additional details

Share job