Our client is an elite, globally integrated financial institution renowned for its market-leading operational stability and continuous investment in modern technology. Operating at a massive scale, they foster a tech-first engineering culture that prioritizes deep technical problem-solving over bureaucratic processes. This organization is highly valued by its employees for its long-term career stability, offering an environment where engineers average over 3+ years of tenure.
Â
The Role & Landscape
You will take complete ownership of regional projects, modernizing monitoring frameworks, and working closely with global leadership to ensure the absolute reliability of critical distributed trading systems.
Build & Modernize: Design, enhance, and implement enterprise-scale Prometheus and Logstash pipelines—handling everything from metric scraping and relabeling to advanced log parsing and routing into Elasticsearch.
In-House Tooling:Â Write clean, production-grade automation code (primarily in Python, Go, or Shell) to develop internal tools, eliminate manual toil, and optimize pipeline efficiency.
Infrastructure Optimization:Â Troubleshoot and perform deep-dive performance tuning across enterprise Linux (RHEL 7/8/9) server clusters.
Collaborative Impact:Â Act as the technical bridge between local infrastructure teams and global engineering hubs, orchestrating critical incident resolution, DR drills, and long-term capacity planning.
Â
What We Are Looking For
Education: A completed full-time Bachelor’s degree in Computer Science, Electronic Engineering, or a strictly related technical discipline.
Experience: 8–10 years of professional IT experience, with at least 3+ years dedicated to Site Reliability Engineering (SRE) or advanced Observability frameworks within a complex environment (Investment Banking or Financial Services experience is preferred).
The Coding Baseline:Â Strong foundational programming skills. You must be comfortable writing and explaining script logic/tooling in Python, Shell, or Go.
The Stack:Â Proven hands-on expertise building and maintaining pipelines across Prometheus, Grafana, VictoriaMetrics, and the ELK stack (Logstash/Elasticsearch). Deep comfort with RHEL administration and systems-level fault finding.
Professional Stability:Â A proven track record of career stability, ideally showcasing 3+ years of tenure with previous employers.
Communication:Â Absolute fluency in professional English is mandatory; conversational proficiency in Chinese (Cantonese or Mandarin) is highly advantageous for regional team collaboration.
Â
If this outstanding opportunity sounds like your next career move, please submit through "Apply Now" or send your resume in Word format to Lu Zhang at resume@pinpointasia.com and put Lead SRE & Observability Engineer - Global Financial Firm - J12933 in the subject header.
Â
Data provided is for recruitment purposes only.
