IT Security Engineering Associate Director
IT · Full-time
Hyderabad, Telangana, India
Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.
The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.
Pay and Benefits:
- Competitive compensation, including base pay and annual incentive
- Comprehensive health and life insurance and well-being benefits, based on location
- Pension / Retirement benefits
- Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
- DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).
We are seeking an experienced Associate Director, IAM Site Reliability Engineering (SRE), Observability & Automation to lead reliability engineering, operational architecture, observability, service automation, and platform resiliency initiatives across the enterprise Identity and Access Management (IAM) ecosystem.
This role is responsible for improving the availability, performance, scalability, recoverability, and operational readiness of critical IAM platforms, including Privileged Access Management (PAM), Active Directory, PKI, Secrets Management, Authentication Services, Cloud Identity Platforms, and Identity Security services.
The ideal candidate combines deep expertise in IAM operations, site reliability engineering, observability, scripting, and test automation to drive proactive monitoring, service health management, operational automation, and continuous improvement across the IAM landscape.
The Associate Director will work closely with Engineering, Security, Infrastructure, and Architecture teams to ensure IAM services are designed, instrumented, tested, and operated for maximum reliability and resilience.
Key Responsibilities
IAM Reliability Engineering & Operations
- Lead reliability engineering initiatives across IAM platforms and services.
- Drive platform availability, resiliency, service health, and operational readiness improvements.
- Establish and implement SRE best practices and operational excellence standards.
- Partner with engineering teams to embed reliability, monitoring, and automation requirements throughout the development lifecycle.
- Perform production readiness reviews and operational risk assessments.
Observability & Monitoring Engineering
- Design and implement monitoring and observability strategies for IAM platforms.
- Develop standards for:
- Infrastructure Monitoring
- Application Monitoring
- User Experience Monitoring
- Transaction Monitoring
- Dependency Monitoring
- Security Event Monitoring
- Cloud Service Monitoring
- Establish logging, metrics, tracing, telemetry, dashboards, and alerting standards.
- Build and manage centralized observability solutions leveraging platforms such as Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, and Application Insights.
- Develop service health dashboards and executive operational reporting.
Reliability Architecture & Platform Engineering
- Create operational architecture diagrams, service dependency maps, data flow diagrams, and resiliency models.
- Review platform designs and identify scalability, performance, availability, and resiliency risks.
- Define and implement standards for:
- High Availability (HA)
- Disaster Recovery (DR)
- Failover Design
- Capacity Planning
- Fault Tolerance
- Service Recovery
- Collaborate with architecture and engineering teams to eliminate single points of failure.
Automation Engineering & Testing
- Develop automation frameworks to improve operational efficiency and platform reliability.
- Create scripts and tooling using PowerShell, Python, Bash, or similar technologies for monitoring, remediation, health validation, and operational support.
- Design and implement automated testing frameworks for IAM platforms, APIs, integrations, and provisioning workflows.
- Develop automation for:
- Platform Health Checks
- Service Validation
- Monitoring Configuration
- Incident Triage
- Alert Enrichment
- Automated Recovery Tasks
- Operational Reporting
- Drive Infrastructure as Code (IaC) and configuration automation practices.
- Partner with engineering teams to integrate automated testing into CI/CD pipelines.
Availability & Resilience Management
- Define and monitor SLIs, SLOs, Error Budgets, and Availability Targets.
- Drive initiatives to improve uptime, performance, reliability, and recoverability.
- Lead DR testing, failover exercises, resiliency validation, and business continuity readiness activities.
- Identify operational bottlenecks and implement preventive controls.
Incident & Service Health Management
- Lead major incident response, technical troubleshooting, and root cause analysis.
- Drive post-incident reviews and corrective action programs.
- Analyze incident trends and recurring service issues.
- Implement proactive monitoring and predictive alerting solutions.
- Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
Required Qualifications
- 10+ years of experience in IAM, Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or Cybersecurity.
- Strong experience supporting enterprise IAM technologies, including:
- PAM
- Active Directory
- PKI
- Authentication Services
- Secrets Management
- Azure AD / Entra ID
- Cloud Identity Platforms
- Expertise in observability, monitoring design, telemetry, and service instrumentation.
- Hands-on experience with enterprise monitoring platforms including Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, or equivalent.
- Strong scripting and automation experience using PowerShell, Python, Bash, or similar technologies.
- Experience building automated testing frameworks and validation solutions for enterprise platforms.
- Experience with APIs, workflow automation, CI/CD pipelines, and Infrastructure as Code.
- Strong understanding of High Availability, Disaster Recovery, Resilience Engineering, and Business Continuity.
- Experience conducting root cause analysis and driving operational improvement initiatives.
- Strong communication, stakeholder management, and leadership skills.
We are seeking an experienced Associate Director, IAM Site Reliability Engineering (SRE), Observability & Automation to lead reliability engineering, operational architecture, observability, service automation, and platform resiliency initiatives across the enterprise Identity and Access Management (IAM) ecosystem. This role is responsible for improving the availability, performance, scalability, recoverability, and operational readiness of critical IAM platforms, including Privileged Access Management (PAM), Active Directory, PKI, Secrets Management, Authentication Services, Cloud Identity Platforms, and Identity Security services. The ideal candidate combines deep expertise in IAM operations, site reliability engineering, observability, scripting, and test automation to drive proactive monitoring, service health management, operational automation, and continuous improvement across the IAM landscape.