Role Summary
This is a senior, client-facing onshore role requiring deep technical expertise, strong leadership presence, and the ability to independently drive architecture decisions across large, complex programs while managing distributed (onshore/offshore) engineering teams. A very strong expertise in working on Life Sciences Clinical R&D specifically Veeva CTMS is an important aspect for this role.
Key Responsibilities
- Own and drive the enterprise data architecture roadmap across Azure, aligning with business strategy and long-term scalability goals.
- Architect complex, high-volume data platforms spanning batch, real-time, and streaming pipelines using Azure Synapse Analytics, Azure Data Factory, Azure Databricks, Event Hubs/Stream Analytics, and Data Lake Storage Gen2.
- Lead enterprise data modeling efforts (conceptual, logical, physical) across multiple domains — supporting analytics, BI, and AI/ML use cases.
- Define and enforce data governance, lineage, cataloging, and quality frameworks (Azure Purview or equivalent).
- Own the security and compliance architecture — encryption, RBAC, Private Endpoints, Key Vault, data masking, and regulatory alignment (GDPR, HIPAA, GxP).
- Serve as the primary technical authority in client architecture reviews, steering committees, and executive presentations.
- Lead and mentor a team of data architects and senior data engineers, both onshore and offshore.
- Drive modernization/migration initiatives — legacy on-prem warehouses to Azure, or multi-cloud to Azure consolidation.
- Establish and govern CI/CD, IaC (Terraform/Bicep), and DevOps practices for data platform delivery.
- Own cost optimization and FinOps practices across Azure data services.
- Conduct architecture governance reviews, risk assessments, and design authority sign-offs (HLD/LLD).
- Partner with enterprise architects, security teams, and business leadership to ensure alignment with overall IT strategy.
- Evaluate and pilot emerging Azure data/AI capabilities (e.g., Microsoft Fabric, AI Foundry integration) and recommend adoption.
Required Skills & Experience
- 10+ years in data engineering/architecture, with hands-on experience in Azure data services at an architect level.
- Deep, demonstrable expertise across:
- – Azure Data Factory / Synapse Pipelines
- – Azure Synapse Analytics (dedicated & serverless SQL pools)
- – Azure Databricks (PySpark, Delta Lake, Unity Catalog)
- – Azure Data Lake Storage Gen2
- – Azure SQL Database / Managed Instance / Cosmos DB
- – Microsoft Fabric (working knowledge strongly preferred)
- At least one implementation of a Life Sciences Clinical R&D project, either EDC, CTMS or eTMF implementation is required. Veeva CTMS implementation is preferred.
- Proven experience architecting Lakehouse, Medallion architecture, and modern data mesh patterns.
- Strong command of data modeling (3NF, Star/Snowflake, Data Vault 2.0).
- Advanced proficiency in SQL, Python/PySpark, and Spark performance tuning.
- Deep knowledge of Azure security & networking (Managed Identity, Private Link, VNet integration, Azure Firewall).
- Hands-on experience with Infrastructure as Code (Terraform or Bicep) and Azure DevOps CI/CD.
- Strong background in MDM, data governance, and metadata management tools (Purview, Collibra, or similar).
- Track record of leading large-scale migrations (Teradata/Netezza/On-prem SQL Azure).
Excellent stakeholder management — able to independently lead C-level/VP-level architecture