Job Description
The main function of the Data Engineer is to develop, evaluate, test, and maintain architectures and data solutions within our organization. The typical Data Engineer executes plans, policies, and practices that control, protect, deliver, and enhance the value of the organization s data assets.
Day-to-day responsibilities
- Build and maintain data mitigation pipelines that scan, classify, and remediate research datasets before research use.
- Create and manage Hive tables and namespaces for mitigated outputs.
- Move data across Manifold, S3, and NFS
- Debug and triage pipeline failures (workflow conflicts, rate limits, cross-namespace query errors) and document runbooks.
- Request, track, and validate dataset read/write access grants for mitigation workflows.
- Author and iterate on agentic skills/automation (e.g., fdp-mitigate) to make mitigation repeatable and self-service.
- Partner with AI researchers and XFN reviewers on dataset readiness, and report status to the data engineering lead
Key Projects
AI dataset mitigations (e.g. CSAM image/video mitigation. copyright) using agentic tooling. MAANG Experience Highly Wanting.
Must-Have Skills
- Strong Python, SQL/Presto for large-scale batch data processing.
- hands-on data pipeline and workflow engineering orchestration, restart ability, idempotency, debugging at scale.
- practical data-quality and validation rigor (schema checks, dedupe, threshold gating) with clear documentation habits.
Nice-to-Have Skills
- Media/video data processing decode, clipping, frame extraction, embeddings.
- Experience with AI coding agents / prompt-and-skill authoring to automate repetitive data workflows.
Years of Experience: 04 - 05 Years
Education
BS in Computer Science, Data Engineering, or a related technical field; equivalent hands-on data pipeline experience accepted in lieu of a degree. Advanced degree not required.
Interviews: 02 Rounds | Technical / Behavioral | 30 - 45 Minutes.