Role Summary
Support engineer will provide L2/L3 production support for business-critical applications built on Appian, running on AWS, and integrated with enterprise data platforms. The role covers incident, problem, change, and release management, proactive monitoring/observability, operational automation, and continuous service improvements while adhering to SLA/OLA and security/compliance (e.g., GDPR) standards.
Key Responsibilities: Incident & Problem Management
- 5–8 years total experience in application/production support, with 3+ years across Appian and AWS, and strong SQL/Data skills.
- Own and resolve L2/L3 incidents, service requests, and problems within SLA; drive root cause analysis (RCA) and implement permanent fixes.
- Create/run runbooks, SOPs, and knowledge articles; reduce repeat incidents via automation and preventive actions.
- Participate in on-call rotations and major incident bridges; communicate updates in German to business stakeholders.
- Monitor Appian Health Check, system logs, and resource utilization (engines, data servers, JMS, sites, nodes).
- Deploy Appian packages (versions, patches, hotfixes) across environments; validate design-time and runtime issues.
- Troubleshoot process models, integrations (APIs, web services, RPA), records, task queues, and IDP/document management flows.
- Work with developers to fix performance bottlenecks, stuck instances, and data sync issues.
- Operate and troubleshoot workloads on EC2, RDS/Aurora, S3, EFS, ELB/ALB, CloudWatch, CloudTrail, SNS/SQS, Lambda, VPC/IAM, Secrets Manager.
- Implement/monitor logging & observability (CloudWatch metrics/alarms, log insights, dashboards), backup/restore, patching, cost & tag hygiene.
- Support CI/CD pipelines (e.g., Git, Jenkins/GitLab CI, Code Pipeline), infrastructure-as-code (CloudFormation/Terraform – nice to have).
- Diagnose and optimize SQL (preferably PostgreSQL/MySQL/Aurora); query tuning, indexing, explain plans.
- Support data feeds/ETL (e.g., Glue/Athena or on-prem tools), batch jobs, API integrations, data quality checks, and recon.
- Ensure data privacy/compliance (GDPR), secure handling of PII, and proper access controls.
- Plan/execute changes with rollback strategies; maintain configuration baselines and environment parity.
- Track and report SLA/KPI compliance: MTTR, MTTD, incident volume, change success rate, problem elimination, availability.
- Partner with Dev/QA/Platform teams to prioritize stability, performance, scalability, and cost optimization.
- Deliver clear German status updates to business users and leadership during incidents and planned changes.
- Document post-incident reviews (PIR), RCAs, and improvement actions; maintain transparent backlog and timelines.
- 5+ years in application/production support with L2/L3 ownership.
- Hands-on Appian support (health checks, deployments, logs, process/integration troubleshooting).
- AWS operations proficiency (EC2, RDS, S3, VPC, IAM, CloudWatch, backups, patching, alarms).
- Strong SQL and data troubleshooting; familiarity with ETL/data pipelines and API integrations.
- Experience with ITIL processes (Incident/Problem/Change), on-call, and SLA-driven environments.
- Tools: ServiceNow/Jira, Confluence, Git, Jenkins/GitLab, Postman, Splunk/ELK/CloudWatch Logs.
- German B2/C1 and English communication skills.
- Appian Certified Associate/Designer (or higher).
- AWS Certified Solutions Architect – Associate (or similar).
- Scripting/automation with Python or Bash; Terraform/CloudFormation basics.
- Exposure to Appian Cloud operations, Kubernetes/EKS, Prometheus/Grafana, or DataOps.
- Knowledge of security controls (network/firewall/IAM), SSO/SAML/OAuth, and vulnerability remediation.
- Bachelor’s in computer science, Information Technology, or equivalent experience.
- Shift window: EU business hours with flexibility for releases; on-call rotation for P1/P2 incidents.
- Occasional travel to customer sites (as required).
- MTTR reduction and SLA adherence.
- Decrease in repeat incidents via problem management.
- Change success rate and zero-defect releases.
- Improved system performance and cost efficiency on AWS.
- Positive stakeholder satisfaction (CSAT/NPS).