HelloKindred Logo

HelloKindred

SaaS Monitoring Engineer

Posted 14 Days Ago
Be an Early Applicant
Hybrid
London, England, GBR
Mid level
Hybrid
London, England, GBR
Mid level
Design, implement, and maintain monitoring and observability solutions for SaaS platforms and cloud infrastructure. Build centralized dashboards, integrate logs/metrics/traces, configure alerting and automation, investigate incidents, refine thresholds to reduce noise, and collaborate with DevOps and engineering teams to improve reliability and operational excellence.
The summary above was generated by AI
Company Description

Who is HelloKindred?

HelloKindred are specialists in staffing marketing, creative and technology roles, offering a range of talent solutions that can be delivered on-site, remotely or hybrid.

Our vision is to make work accessible and people’s lives better. We do this by disrupting traditional employment barriers – connecting ambitious talent to flexible opportunities with trusted brands.

Job Description

Anticipated Contract End Date/Length: December 18, 2026
Work Set Up: Hybrid (60% Office - 40% Home)
Clearance Required: BPSS Eligibility Required

Our client in the Information Technology and Services industry is looking for a SaaS Monitoring Engineer to join a growing cloud operations team. This role is responsible for designing, implementing, and maintaining monitoring solutions that ensure the health, performance, availability, and reliability of Software-as-a-Service (SaaS) platforms. The successful candidate will play a key role in proactive issue detection, operational excellence, incident management, and observability initiatives, while developing centralized monitoring dashboards that provide real-time visibility into service health and performance across distributed cloud environments.

Responsibilities:

  • Design, implement, and maintain monitoring frameworks for SaaS applications and cloud infrastructure.
  • Monitor system health, availability, latency, error rates, resource utilization, and overall platform performance across distributed environments.
  • Enhance observability through the implementation and optimization of logs, metrics, and traces using modern monitoring platforms.
  • Develop and maintain a centralized console dashboard that provides a real-time view of SaaS service health and operational metrics.
  • Configure dashboard views to deliver actionable insights on service uptime, API performance, incident alerts, dependency status, and business-critical metrics.
  • Integrate data from multiple monitoring and operational sources into unified visualization platforms.
  • Optimize dashboard usability and reporting capabilities for engineering, operations, and leadership stakeholders.
  • Establish intelligent alerting mechanisms to detect anomalies, performance degradation, and service disruptions.
  • Investigate incidents, identify root causes, and implement corrective and preventive measures.
  • Collaborate with DevOps and engineering teams during incident response activities and post-incident reviews.
  • Automate monitoring processes, alert escalation workflows, and operational response procedures.
  • Refine alert thresholds and monitoring configurations to reduce noise and improve signal accuracy.
  • Implement predictive monitoring techniques to proactively identify potential service issues and outages.
  • Partner with software engineering, DevOps, and product teams to incorporate monitoring requirements throughout the development lifecycle.
  • Analyze SaaS performance trends and provide reporting on operational risks and service reliability.
  • Document monitoring strategies, configurations, standards, and best practices.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Proven experience monitoring cloud-based applications, SaaS platforms, or distributed systems.
  • Strong understanding of distributed systems architecture, microservices, and cloud platforms including AWS, Azure, or GCP.
  • Hands-on experience with monitoring, observability, and visualization tools such as Grafana, Prometheus, ELK Stack, Splunk, Datadog, Azure Monitor, or similar technologies.
  • Experience designing and developing interactive dashboards and console-based monitoring solutions.
  • Proficiency in scripting or programming languages such as Python, Bash, or Go.
  • Familiarity with containerization and orchestration technologies including Docker and Kubernetes.
  • Strong analytical, troubleshooting, and problem-solving capabilities.
  • Demonstrated ability to manage incidents, identify root causes, and implement sustainable solutions.
  • Excellent verbal and written communication skills with the ability to translate technical findings into actionable business insights.
  • Experience working in collaborative, fast-paced technology environments.
  • Knowledge of Site Reliability Engineering (SRE) principles and practices preferred.
  • Familiarity with CI/CD pipelines and DevOps methodologies preferred.
  • Exposure to AIOps solutions or machine learning-based monitoring technologies preferred.
  • Cloud platform or monitoring technology certifications preferred.
  • Ability to maintain a proactive approach to operational excellence and continuous improvement.
  • Eligibility to obtain and maintain BPSS clearance.

Additional Information

Please submit a CV/resume (mandatory) along with your application.

All your information will be kept confidential according to EEO guidelines.

Candidates must be legally authorized to live and work in the country where the position is based, without requiring employer sponsorship.

HelloKindred is committed to fair, transparent, and inclusive hiring practices. We assess candidates based on skills, experience, and role-related requirements.

We appreciate your interest in this opportunity. While we review every application carefully, only candidates selected for an interview will be contacted.

HelloKindred is an equal opportunity employer. We welcome applicants of all backgrounds and do not discriminate on the basis of race, colour, religion, sex, gender identity or expression, sexual orientation, age, national origin, disability, veteran status, or any other protected characteristic under applicable law.

Similar Jobs

7 Minutes Ago
Hybrid
London, England, GBR
Senior level
Senior level
Fintech • Mobile • Payments • Software • Financial Services
Lead a global People Technology engineering team to own end-to-end employee journeys, deliver the technical roadmap, improve HRIS integrations and reliability, support incidents and SLAs, standardise agile rituals, and coach engineers to improve adoption and internal customer experience.
Top Skills: APIsAutomated WorkflowsCore Hris SystemsData MappingHrisJira Product DiscoveryJira SoftwareOnboarding SystemsRecruitment ToolsSaaS
51 Minutes Ago
Easy Apply
Hybrid
London, Greater London, England, GBR
Easy Apply
Mid level
Mid level
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Manage end-to-end presence at third-party events across EMEA: strategy, logistics, onsite execution, staffing, vendor management, sponsorship evaluation, lead capture, KPI tracking, and reporting. Improve event workflows, leverage AI tools for research and reporting, partner with EMEA teams, and travel up to 50%.
Top Skills: ChatgptClaudeGeminiMetabaseSalesforceTableau
An Hour Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Integrate autonomy software onto unmanned aircraft, prepare flight-ready systems, support on-site flight tests, debug hardware/software issues, capture and analyze flight data, collaborate with autonomy/GNC/systems teams, and support certification and continuous improvement for mission deployments.
Top Skills: AvionicsC++DdsEmbedded SystemsGncLinuxPythonRosRtosSensor FusionShell ScriptingSimulation Tools

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account