Incident Management
bounteous · London
Experience: 8+ years of experience
Incident Manager/Commander with SOC 1 or SOC 2 Audit experience
Job Description – Incident Commander
UK - Remote
Department: Service Management / Technology Operations.
About the Role
We are seeking a highly skilled Incident Commander to join our Fintech operations team. This role is critical in ensuring the stability, resilience, and continuous availability of our trading and financial platforms. The Incident Commander will lead the management of high- everity incidents, drive root cause analysis, enforce SLAs, and ensure operational readiness through automation, disaster recovery, and business continuity planning.
You will be responsible for end-to end Incident & Problem Management, overseeing command during outages, and partnering closely with engineering, infrastructure, and business stakeholders.
Key Responsibilities
Incident & Problem Management
• Lead and manage IT incidents, including identification, triage, resolution, documentation, and post-incident reviews for critical and high-severity issues.
• Serve as the primary Incident Commander during major outages, ensuring swift coordination between technical teams and business stakeholders.
• Conduct incident washups to identify learnings, drive accountability, and prevent repeat issues.
• Manage and monitor SLAs, ensuring timely resolution and proper lifecycle governance.
• Own Problem Management lead root cause analysis (RCA), track problem records, and oversee corrective/preventive actions.
Service Management & Reporting
• Provide downtime and availability reporting related to application outages.
• Generate monthly Service Management Risk Reports, highlighting key operational risks and remediation status.
• Ensure effective Service Management reporting aligned with the Service Catalog and business expectations (availability, efficiency, ticket handling, customer satisfaction).
• Oversee data integrity across ITSM repositories/tools (e.g., ServiceNow, JSM, HPSM, CMDB).
Audit, Compliance & Process Governance
• Own and maintain audit-ready documentation for operational processes, policies, runbooks, standard operating procedures (SOPs), and control frameworks (e.g., Risk Control Matrices).
• Actively participate in internal and external audit cycles, including SOC Type 1 & Type 2 audits and regulatory examinations (e.g., CFTC, where applicable).
• Join audit calls and walkthroughs with auditors, providing evidence, clarifications, and demonstrations of control effectiveness.
• Ensure all incident, problem, and change management processes are documented, version-controlled, and aligned with audit and regulatory requirements.
• Drive process management by defining, reviewing, and continuously improving operational workflows to meet compliance standards.
• Prepare and deliver audit-related reports, including control testing results, gap analyses, remediation trackers, and compliance dashboards.
• Coordinate with cross functional teams (Operations, Engineering, Risk, Compliance) to ensure timely closure of audit findings and corrective actions.
• Maintain a centralized repository of all audit artifacts, ensuring accessibility and completeness for scheduled and ad-hoc audit requests.
Governance & Continuous Improvement
• Run weekly Release Washups and participate in Change Advisory Board (CAB) meetings to ensure safe and reliable deployments.
• Contribute to building and maintaining a comprehensive CMDB.
• Drive harmonized adoption of the Service Management framework across teams.
• Support automation initiatives to reduce MTTR, improve monitoring/alerting, and streamline incident handling.
Resilience & Continuity
• Develop and manage Disaster Recovery (DR) Plans; lead DR testing exercises.
• Operationalize the Business Continuity Plan (BCP) across IT and business units.
Qualifications & Skills
Must Have
• 8+ years of experience in Incident Management, Problem Management, and IT Operations within fintech, banking, or trading environments.
• Proven experience as Major Incident Manager/Incident Commander handling real-time high-severity events.
• Strong knowledge of ITIL processes (Incident, Problem, Change, Knowledge, Configuration Management).
• Hands-on experience with ITSM tools (ServiceNow, JSM, HPSM, CMDB).
• Experience in reporting & analytics (availability, SLA compliance, RCA trends).
• Experience in audit support, process documentation, and working with internal/external auditors.
• Excellent communication and stakeholder management skills; ability to lead under pressure.
• Good understanding of high-level architecture in software as well as infrastructure.
Nice to Have
• Knowledge of trading systems, order management platforms, or exchanges.
• Experience in running Splunk queries, understanding basic troubleshooting.
• ITIL v3/v4 certification or equivalent.
• Prior experience in global 24×7 financial services environments.
• Familiarity with SOC audit frameworks and regulatory compliance in financial services.
• Prefers working in night shift
Job Description – Incident Commander
UK - Remote
Department: Service Management / Technology Operations.
About the Role
We are seeking a highly skilled Incident Commander to join our Fintech operations team. This role is critical in ensuring the stability, resilience, and continuous availability of our trading and financial platforms. The Incident Commander will lead the management of high- everity incidents, drive root cause analysis, enforce SLAs, and ensure operational readiness through automation, disaster recovery, and business continuity planning.
You will be responsible for end-to end Incident & Problem Management, overseeing command during outages, and partnering closely with engineering, infrastructure, and business stakeholders.
Key Responsibilities
Incident & Problem Management
• Lead and manage IT incidents, including identification, triage, resolution, documentation, and post-incident reviews for critical and high-severity issues.
• Serve as the primary Incident Commander during major outages, ensuring swift coordination between technical teams and business stakeholders.
• Conduct incident washups to identify learnings, drive accountability, and prevent repeat issues.
• Manage and monitor SLAs, ensuring timely resolution and proper lifecycle governance.
• Own Problem Management lead root cause analysis (RCA), track problem records, and oversee corrective/preventive actions.
Service Management & Reporting
• Provide downtime and availability reporting related to application outages.
• Generate monthly Service Management Risk Reports, highlighting key operational risks and remediation status.
• Ensure effective Service Management reporting aligned with the Service Catalog and business expectations (availability, efficiency, ticket handling, customer satisfaction).
• Oversee data integrity across ITSM repositories/tools (e.g., ServiceNow, JSM, HPSM, CMDB).
Audit, Compliance & Process Governance
• Own and maintain audit-ready documentation for operational processes, policies, runbooks, standard operating procedures (SOPs), and control frameworks (e.g., Risk Control Matrices).
• Actively participate in internal and external audit cycles, including SOC Type 1 & Type 2 audits and regulatory examinations (e.g., CFTC, where applicable).
• Join audit calls and walkthroughs with auditors, providing evidence, clarifications, and demonstrations of control effectiveness.
• Ensure all incident, problem, and change management processes are documented, version-controlled, and aligned with audit and regulatory requirements.
• Drive process management by defining, reviewing, and continuously improving operational workflows to meet compliance standards.
• Prepare and deliver audit-related reports, including control testing results, gap analyses, remediation trackers, and compliance dashboards.
• Coordinate with cross functional teams (Operations, Engineering, Risk, Compliance) to ensure timely closure of audit findings and corrective actions.
• Maintain a centralized repository of all audit artifacts, ensuring accessibility and completeness for scheduled and ad-hoc audit requests.
Governance & Continuous Improvement
• Run weekly Release Washups and participate in Change Advisory Board (CAB) meetings to ensure safe and reliable deployments.
• Contribute to building and maintaining a comprehensive CMDB.
• Drive harmonized adoption of the Service Management framework across teams.
• Support automation initiatives to reduce MTTR, improve monitoring/alerting, and streamline incident handling.
Resilience & Continuity
• Develop and manage Disaster Recovery (DR) Plans; lead DR testing exercises.
• Operationalize the Business Continuity Plan (BCP) across IT and business units.
Qualifications & Skills
Must Have
• 8+ years of experience in Incident Management, Problem Management, and IT Operations within fintech, banking, or trading environments.
• Proven experience as Major Incident Manager/Incident Commander handling real-time high-severity events.
• Strong knowledge of ITIL processes (Incident, Problem, Change, Knowledge, Configuration Management).
• Hands-on experience with ITSM tools (ServiceNow, JSM, HPSM, CMDB).
• Experience in reporting & analytics (availability, SLA compliance, RCA trends).
• Experience in audit support, process documentation, and working with internal/external auditors.
• Excellent communication and stakeholder management skills; ability to lead under pressure.
• Good understanding of high-level architecture in software as well as infrastructure.
Nice to Have
• Knowledge of trading systems, order management platforms, or exchanges.
• Experience in running Splunk queries, understanding basic troubleshooting.
• ITIL v3/v4 certification or equivalent.
• Prior experience in global 24×7 financial services environments.
• Familiarity with SOC audit frameworks and regulatory compliance in financial services.
• Prefers working in night shift