Frequently Asked Question

Important Diagnostics Process - Serverside Systems and Virtualisation
Last Updated 59 minutes ago

For Enterprise Server Systems and Virtualisation GEN operates a comprehensive and detailed Diagnostic processes (SDP)

A structured diagnostic process for server-side systems and virtualisation reduces risk, improves accuracy, and creates a repeatable method for recovery. The model is:

  1. Snapshot
  2. Diagnose
  3. Planning

This approach is designed to avoid hasty changes on live systems and to ensure remedial work is based on evidence rather than guesswork.

1. Snapshot

The first step is to capture the current state of the affected systems and services before making changes.

This is not limited to virtual machine snapshots. It refers more broadly to taking a point-in-time evidence capture of the environment, which may include:

  • Operating system state
  • Running services
  • Process information
  • Network status
  • Storage and filesystem health
  • Cluster or hypervisor state
  • Logs
  • Configuration
  • Performance counters
  • Alerts and error conditions
  • Application-specific diagnostics

In virtualised environments, this may also include:

  • Hypervisor health
  • VM state
  • Host resource utilisation
  • Datastore status
  • Backup status
  • Replication state
  • Snapshot trees
  • HA or cluster membership state

Purpose of the snapshot stage

The snapshot stage creates a reliable baseline of what was happening at the time of the fault. This is important because:

  • Live systems change continuously
  • Logs rotate
  • Services restart
  • Resource usage fluctuates
  • Automatic recovery routines may alter evidence
  • Human intervention can unintentionally destroy useful diagnostic context

By capturing evidence first, the environment can be analysed as it was when the issue occurred.

Important principle

The snapshot stage should happen before trial fixes, reboots, service restarts, failovers, or storage interventions wherever practical. Once a change is made, valuable evidence may be lost.

2. Diagnose

Once the state has been captured, the next step is to analyse the evidence, consider the complaint and identify the actual issue or issues.

Common diagnostic outcomes

The evidence may show:

  • A clear root cause
  • Several plausible causes requiring prioritisation
  • A service dependency issue rather than an application fault
  • A storage, network, or authentication issue presenting elsewhere
  • A capacity problem such as disk, memory, IOPS, or CPU exhaustion
  • A configuration drift issue
  • A failed update or incomplete patch state
  • A hardware degradation issue
  • A cluster quorum or replication issue
  • A backup or restore path issue

Why diagnosis should be separate from remedial actions

Separating diagnosis from remediation avoids:

  • Guesswork
  • Repeated disruptive changes
  • Masking the original fault
  • Introducing new faults while chasing old ones
  • Treating symptoms instead of causes
  • Worsening outages through unplanned intervention

A proper diagnostic phase ensures remedial action is justified by evidence.

3. Planning

After the issue has been identified, the final stage is to plan the recovery work carefully before carrying it out.

Planning turns findings into a controlled remedial procedure.

What planning includes

A good remedial plan should define:

  • The confirmed issue or working diagnosis
  • The intended corrective action
  • Why that action is appropriate
  • Dependencies and prerequisites
  • Risks of the change
  • Expected effect on systems and users
  • Rollback options

Outcome of the planning stage

The result should be a considered, sequenced set of actions rather than ad hoc experimentation. This gives engineers a clear route to service recovery while controlling risk.

Benefits of GEN's snapshot, diagnose, planning (SDP) approach

This method offers significant operational and technical advantages.

1. Rollback protection

Where supported, point-in-time captures and pre-change records provide a safer route for recovery.

Benefits include:

  • Ability to reverse a bad change
  • Safer execution of higher-risk remedial actions
  • Reduced fear of decisive action when evidence supports it
  • Faster restoration if an attempted fix introduces problems
  • Better change confidence in virtualised environments

This does not remove the need for proper backups, but it does improve short-term operational safety.

2. Retrospective analysis

A captured diagnostic record allows later review even after the live issue has changed or disappeared.

This is valuable because:

  • Intermittent faults may not be visible later
  • Services may restart automatically
  • Evidence may vanish during recovery
  • Technical review often continues after service restoration

Retrospective analysis supports:

  • Root cause analysis
  • Incident reporting
  • Trend identification
  • Improvement of runbooks
  • Future prevention work

3. Better use of machine learning and automation

Structured snapshots create consistent, high-quality diagnostic datasets.

This has practical value for:

  • Pattern recognition
  • Correlating recurring faults
  • Flagging likely root causes
  • Comparing current incidents to historical incidents
  • Building decision support tools
  • Improving alert triage
  • Identifying early warning indicators

Machine learning is only as useful as the quality and consistency of the input data. A disciplined snapshot process creates evidence that can be indexed, compared, and analysed at scale.

4. Considered remediation instead of trial and error

One of the biggest advantages is the move away from random or impulsive fixes.

Without structure, incident response often becomes:

  • Restart a service
  • Reboot a host
  • Clear a cache
  • Move a VM
  • Retry a job
  • Change several settings at once

This can sometimes restore service, but it often creates new problems or hides the original one, and gives limited justification. 

A planned approach provides:

  • Higher confidence in each action
  • Fewer unnecessary changes
  • Lower risk of collateral damage
  • Better identification of what actually fixed the issue
  • Faster long-term resolution, even if the first few minutes feel slower

In critical systems, deliberate action is usually faster overall than repeated failed interventions.

5. Auditable actions

A structured process naturally creates an audit trail.

This is useful for:

  • Internal review
  • Compliance requirements
  • Change control
  • Security investigations
  • Incident reporting
  • Customer reporting
  • Post-incident Reporting

This improves transparency and reduces ambiguity.

6. Accountability

The process supports clear engineering ownership.

Because the work is captured in stages, it is easier to determine:

  • Who collected the evidence
  • Who diagnosed the issue
  • Who approved the plan
  • Who carried out the changes
  • Whether the actions matched the evidence
  • Whether the recovery followed process

This is not about blame. It is about clarity, responsibility, and professional standards in operational work.

7. Training and knowledge transfer

Captured evidence and planned remediation make excellent training material.

Benefits include:

  • Junior engineers can learn from real incidents
  • Senior engineering reasoning becomes visible and teachable
  • Teams can review decision quality, not just outcomes
  • Repeated incident patterns can be turned into runbooks
  • Training can be based on real system states rather than theory alone

This improves team maturity over time and makes knowledge less dependent on a small number of experienced individuals.

8. Improved root cause accuracy

By preserving evidence before intervention, the team is more likely to find the true cause rather than the last visible symptom, leading to faster overall resolutions with increased confidence. 


9. Reduced operational risk

Every unplanned change on a degraded system carries risk. Structured planning reduces that risk by forcing thought before action.

This is particularly important where there are:

  • Production workloads
  • Shared storage
  • Cluster dependencies
  • Replication relationships
  • Business-critical applications
  • Compliance-sensitive systems

10. Better communication during incidents

A clear three-step process improves communication with technical teams, management, and customers.

It provides a simple framework:

  • Evidence has been captured
  • The issue has been identified or narrowed down
  • A recovery plan is being prepared or executed

This avoids vague updates and gives stakeholders a clearer view of progress.


This website relies on temporary cookies to function, but no personal data is ever stored in the cookies.
OK
Powered by GEN UK CLEAN GREEN ENERGY

Loading ...