Introduction

Micro Focus builds enterprise software for monitoring and troubleshooting cloud systems across AWS, Azure, and GCP.

As customer environments scaled, alert volume increased, tooling became fragmented, and investigation paths were unclear. Teams struggled to efficiently diagnose and resolve incidents.

This case study shows how redesigning cloud troubleshooting workflows reduced alert noise, clarified investigation paths, and improved customer confidence during high-pressure incidents.

Investigation-Driven Cloud Troubleshooting Workflow

Investigation-Driven Cloud Troubleshooting Workflow

Investigation-Driven Cloud Troubleshooting Workflow

 

Role & Scope

This initiative was led from a UX strategy and product design perspective in close partnership with Product, Engineering, Security, Customer Success, and Cloud Operations.

The work focused on rethinking investigation workflows end-to-end, clarifying page intent during incidents, and validating concepts through customer design partners before broader rollout.

 

Business Context & Constraints

The platform supported complex, multi-cloud environments. As customers scaled infrastructure, alert volume increased significantly. Investigation often required navigating across disconnected dashboards and manually correlating data.

At the same time, the redesign had to work within an established enterprise platform, shared UI foundations, and evolving architecture. Rebuilding the system was not feasible. Improvements needed to enhance investigation without disrupting current usage or introducing technical risk.

Previous Investigation Tool with Adoption Challenges

 

Researching & Defining the Problem

Research centered on how cloud incidents were investigated in the real world.

Customer interviews and workflow analysis revealed a consistent gap: while the platform exposed large volumes of data, it did not support investigation as a structured process. Users were forced to manually piece together alerts, metrics, and topology across multiple screens.

The core issue was not missing data — it was missing guidance.

Missing Transition Between Interactive Report and Service Details

 

Defining the User

A primary persona was created to anchor decisions.

Olivia, the Domain Expert, represents users responsible for diagnosing cloud issues across infrastructure and services. She receives escalations under time pressure and needs immediate clarity on impact, root cause, and next steps.

Design decisions were evaluated through the lens of this persona to avoid overgeneralization and scope drift.

Primary Persona: Domain Expert

 

Journey & Workflow Mapping

End-to-end investigation journeys were mapped from initial signal to resolution.

This work revealed where users stalled, backtracked, or lost context between tools. It became clear that troubleshooting was a multi-stage workflow — not a single alert interaction — and required structural support across the experience.

Event Response Workflow (Domain Expert)

 

The Core Design Challenge

A key constraint emerged around the platform’s UI foundations.

Existing components were optimized for monitoring and reporting, not investigation. Pages lacked clear intent, visual density made prioritization difficult, and a config-based component system limited flexibility.

Rather than rebuilding the foundation, the challenge became designing investigation-focused workflows within those constraints.

Flexibility Challenges in UI Foundations

 

Strengthening the UI Foundation

To better support diagnostic workflows, new modular components were designed to improve hierarchy, clarity, and intent.

This included recommending a shift away from heavily configuration-driven components toward more flexible, code-based structures. The goal was to enable clearer investigation states without destabilizing the broader platform.

Modular Components for Troubleshooting Workflows

 

Researching Existing Solutions

Competitive and internal tools were reviewed to understand how others approached troubleshooting.

Most platforms emphasized surfacing data rather than guiding decision-making. This reinforced an opportunity to differentiate through clarity of flow rather than additional feature density.

Comparative Analysis of Troubleshooting UX

Competitor: Anomaly and Event Visualization

 

Creating a Vision Blueprint (Flows)

Insights from research and mapping informed a blueprint for a more cohesive investigation experience.

The blueprint defined clear entry points into troubleshooting, structured progression from overview to drilldown, and intent-driven page layouts aligned to investigation stages.

Scenario walkthroughs were used to align stakeholders and secure early buy-in.

Vision Blueprint: End-to-End Troubleshooting Flow

 

Technical Feasibility and Phasing

Engineering collaboration began early to validate feasibility and define realistic scope.

Investigation clarity was prioritized over automation. Anomaly detection was removed from the initial phase due to platform readiness and delivery constraints. Phasing allowed improvements to ship without overextending technical resources.

Phase 1 Scope: UI Foundation-Based Implementation

Future Phase: Anomaly Detection

 

Prototyping and Design Partner Feedback

Wireflows and interactive prototypes were tested with design partners and internal stakeholders.

Feedback focused on navigation clarity, information hierarchy, and investigation continuity. Iteration cycles refined cognitive load and clarified next steps at each stage of diagnosis.

Insights from Design Partner Testing

 

Phase 1: Simplified Metric Visibility

Phase 2: Geographic Mapping & Advanced Visualization

 

Sales Prototype and Final Handoff

A high-fidelity prototype was created to support Sales demos and stakeholder alignment.

This prototype helped communicate the value of the redesigned experience, supported internal buy-in, and served as a reference point during development and handoff.

 

Results

The redesigned troubleshooting workflows reduced alert noise and clarified investigation paths.

Customers were able to diagnose issues with greater confidence and take clearer next steps during incidents. Feedback reflected improved usability and stronger trust in the platform.

The work contributed to improved customer retention, including renewals from accounts previously at risk of churn.