Monitoring and Alerting
Description
We monitor all hardware and software components involved in delivering our services, as well as your application (on an optional basis). This document outlines the alerting procedures for Managed Services. Incoming alerts are categorized and prioritized by our Event Management team within the Service Operations Center. Error classifications, response times, and operational procedures are defined in the Service Level Agreement (SLA) established with you; the SLA also specifies the notification channels. We offer a standard baseline alerting process for all customers. This process can be extended by optionally adding Service Management and further customized by adding Server Management Enterprise.
Basic Process
For every service, we implement a baseline set of checks; details can be found in the respective service description. For alerting, we aim to strike the right balance between costs caused by overly sensitive alerting and potential consequential costs resulting from service downtime. Our best practice involves categorizing alarms into
- Warnings and
- Critical as well as employing a multi-stage detection process for these states whenever possible. Typically, we perform checks at 5-minute intervals; following a status change, we conduct four additional checks at 1-minute intervals. Any settings that deviate from this standard are specified in the respective service descriptions.
Critical Alarms
Depending on the type, a critical alarm may indicate a reduction in service quality. Since false positives cannot be ruled out—and the issue affecting the service might be transient—we classify the initial occurrence as a "soft" alarm and continue to monitor the service. If the check fails on the fourth later attempt, the alarm transitions to a "hard" state; an alert is then triggered under the SLA, and notifications are sent to the contacts and addresses (email or phone number) stored in our alerting system. Depending on the agreement in place, either you report an incident, or our system reports it to us.
Warnings
A warning is generally triggered before a critical alarm and is intended to support an early response. However, a warning is not a confirmed incident; wherever possible, we set thresholds so that no impairment to the affected service is observed at the warning stage. It indicates that a critical alarm may be triggered soon. We handle warnings on a "best-effort" basis, analyzing data trends—for example—to prevent impending incidents whenever possible. The distinction between soft and hard states applies to warnings as well.
Extended Process
This extension allows you to customize the thresholds for triggering warnings and critical alarms to suit the specific needs of your setup.
Customized Process
Customers with Service Management Enterprise can define thresholds and the transitions between "soft" and "hard" alarm states for each alarm. Additionally, an optional incident response manual can be created to establish a customized framework for resolving error states.
Our goals for this service
- Prevent impending service disruptions
- Quickly notify the appropriate contact person after a disruption
What you can expect from us
- Basic checks based on our best practices
- Customized checks (optional)
Optional Service Features
- Direct alerting by us (see Managed Hosting Service Level Agreement)
- Additional checks
- Extended process
- Customized process
Service Requests
Contact our support and use the following lines as the template to submit a standard service request:
- Change contact details for receiving alarm notifications
- Set up a customized check
- Change check intervals (Service Management Enterprise only)
Important Information
- The contact details you entered in the customer frontend are not used by our alarm system; they are intended only for communication by our staff.
- Please ensure that the contact information in our alarm system is always up to date. You are welcome to submit a ticket via our support for this purpose.