All Jobs Vacancy

Data Scientist 3 (Big data Engineer) || Remote

Posted 1 day ago by NJTECH INC.

Responsibilities

  • Implement ETL/ELT workflows for both structured and unstructured data
  • Collaborate with cross-functional teams including data scientists, analysts, and stakeholders
  • Design and maintain data models, schemas, and database structures to support analytical and operational use cases
  • Implement data validation and quality checks to ensure accuracy and consistency
  • Contribute to data governance initiatives, including metadata management, data lineage, and data cataloging
  • Implement data security measures, including encryption, access controls, and auditing; ensure compliance with regulations and best practices
  • Working in agile, multicultural environments
  • Automate deployments using CI/CD tools
  • Evaluate and implement appropriate data storage solutions, including data lakes (Azure Data Lake Storage) and data warehouses
  • Proficiency in Python and R programming languages
  • Strong SQL querying and data manipulation skills
  • Experience with Azure cloud platform
  • Experience with DevOps, CI/CD pipelines, and version control systems
  • Strong troubleshooting and debugging capabilities
  • Design and develop scalable data pipelines using Apache Spark on Databricks
  • Optimize Spark jobs for performance and cost efficiency
  • Integrate Databricks solutions with cloud services (Azure Data Factory)
  • Ensure data quality, governance, and security using Unity Catalog or Delta Lake
  • Deep understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL
  • Hands-on experience with Databricks notebooks, clusters, jobs, and Delta Lake
  • Build AI accelerators and reusable components to boost development teams productivity.
  • Implement Gen AI / LLM application and utilities (Agentic AI, Harness / Prompt engineering, RAG)
  • Implement AI application governance, observability and evaluation framework
  • Design & Implement vector stores for knowledge bases and agent memory
  • Integrate AI applications / utils / tools with enterprise traceability and SIEM tools
  • Knowledge of ML libraries (MLflow, Scikit-learn, TensorFlow)
  • Databricks Certified Associate Developer for Apache Spark
  • Azure Data Engineer Associate
Rate:
Not specified
Location:
Remote
IR35 Status:
Outside
Remote Status:
Remote
Industry:
Data & Analytics
Seniority Level:
Not Specified

Take-Home Pay

Not Available

Visit calculators for additional details

Share job