Position Purpose
The Senior Data Scientist / Machine Learning Engineer will lead the design, development, and deployment of end-to-end Document Intelligence solutions for the U.
S. Fish and Wildlife Service (USFWS). In this role, you will build robust document-ingestion, OCR, field-extraction, free-text remediation, and classification pipelines to process permits, certificates, and legacy records. A core focus will be implementing confidence-based human-review workflows and maintaining strict, precise source traceability across all extracted data points.
Key Responsibilities
- Pipeline & OCR Engineering: Build and optimize end-to-end document processing pipelines, combining OCR, layout-aware models, and NLP techniques to extract structured fields and remediate unstructured free text from permits, certificates, and scanned legacy records.
- Traceability & Bounding Box Mapping: Establish precise document-coordinate mapping (bounding boxes/character-level offsets) to map extracted fields back to original source documents for complete auditability.
- Human-in-the-Loop Workflow Design: Design confidence-scoring mechanisms and route low-confidence or high-risk extractions to human-in-the-loop (HITL) review interfaces.
- Model Evaluation & Threshold Tuning: Perform rigourous evaluation on labeled datasets, analyzing precision, recall, and edge cases to calibrate operational confidence thresholds.
- Cloud Architecture & Data Modeling: Implement scalable processing pipelines using AWS cloud services, PostgreSQL/SQL databases, document data models, and RESTful APIs.
- Security & Compliance: Ensure all solutions comply with federal security frameworks (FISMA, FedRAMP, NIST SP 800-53) and Privacy Act controls, designing explainable and auditable automated decisions.
Required Capabilities
- Data Science & ML Expertise: Senior-level experience in applied machine learning/data science focused on OCR, document understanding, information extraction, NLP, or text classification.
- Python & Evaluation Mastery: Advanced Python skills with demonstrated experience in model evaluation, precision/recall metrics, threshold calibration, and structured error analysis.
- Traceable HITL Systems: Proven track record of building traceable, auditable human-review workflows for low-confidence AI predictions.
- Data Systems & Integration: Strong working knowledge of SQL/PostgreSQL, semi-structured document models, API development, and cloud object storage (e.g., S3).
- Governance & Security: Experience designing audit trails for automated decision systems and safeguarding sensitive or PII data.
Preferred Qualifications
- Education: Bachelor’s degree or higher in Computer Science, Data Science, Machine Learning, Computational Linguistics, or a related quantitative field.
- AWS Ecosystem: Hands-on experience with AWS AI/ML and serverless services, such as Textract, Bedrock, SageMaker, Lambda, S3, and DynamoDB (or federal cloud equivalents).
- Government Document Processing: Prior experience extracting structured data from government forms, permits, environmental certificates, scientific publications, or historical scanned archive collections.
- Advanced Document Understanding: Experience with layout-aware vision-language models, image preprocessing (deskewing, binarization, noise reduction), taxonomy alignment, entity linking, and MLOps practices.
- Federal Compliance & Controls: Familiarity with FISMA, NIST SP 800-53, FedRAMP compliance, the Privacy Act, model explainability, and AI model risk management principles.