AI Data Knowledge Engineer
- Hybrid- Remote
- 6 months +
- Inside IR35
- SC Clearance (or eligible)
Our customer, a central government organization, are building an enterprise-grade AI platform to enable secure, scalable and production-ready use of Generative AI and ML across the programme.
The successful candidate will have experience in:
- Building integrations against source-system APIs (webhooks, REST, event streams) - connecting CDS estate systems to the platform
- Event-driven ingestion patterns - API Gateway/Lambda receivers, EventBridge, SQS per-consumer buffering
- State/watermark tracking (DynamoDB) and reconciliation to catch delivery gaps
- Python for pipeline/integration development
- Data migration experience ( eg AWS DMS) for onboarding new sources
Nice to have:
- Experience integrating with Confluence, JIRA, or similar enterprise collaboration/ITSM tool APIs
- Apache Airflow and/or dbt
- Enterprise security and governance awareness (regulated environments)
The Data Engineer is responsible for integrating CDS source systems into the AI platform - building the connectors and ingestion pipelines that get source-system data flowing into the platform reliably.
Key Responsibilities:
- Build source-system integrations (webhooks API Gateway/Lambda EventBridge SQS)
- Implement watermark/state tracking (DynamoDB) and periodic reconciliation to catch webhook delivery gaps without placing polling load on source systems
- Onboard new CDS source systems onto the platform's ingestion pattern as scope expands ( eg Confluence, JIRA, and others across the CDS estate)
- Handle data migration activities where a source system needs backfilling into the platform ( eg AWS DMS)
- Work with Knowledge Engineers to ensure ingested data meets the quality/structure needed for downstream graph population
- Maintain integration reliability - monitoring, error handling, and alerting on ingestion failures