IT Infrastructure Stabilization for a Mid-Market Manufacturer
At a glance
- 65% reduction in unplanned downtime within two quarters
- Average helpdesk response cut from four hours to under 30 minutes
- Zero recurring server incidents after hardware replacement
- 100% SLA compliance following rollout
A mid-market manufacturer running two production sites was losing production time to server failures that nobody could predict and helpdesk delays that left shop floor faults sitting for most of a shift. 12th Wonder established continuous infrastructure monitoring, restructured the helpdesk onto a tiered SLA model, traced the recurring failures to their actual source and replaced the hardware behind them. The deployment ran eight weeks and transitioned into an ongoing managed service.
About the client
The client is a mid-market manufacturer operating two production sites, with an internal IT team responsible for both daily user support and the infrastructure those sites depend on.
Its technology environment supports production systems, shop floor devices and back-office users across both locations. When infrastructure stops, production stops with it, which makes availability an operational concern rather than an IT one.
The internal team was small and stretched across two competing demands. Support tickets arrived continuously and infrastructure upkeep had to fit around them, which meant the work that would have prevented outages was consistently displaced by the work created when outages happened.
The challenge
Servers were failing without warning several times a month. Each outage pulled production lines to a stop while the team traced the cause after the fact, which meant every failure cost production time twice, once during the outage and again during the investigation.
Helpdesk's response averaged four hours. A login issue or a printer fault on the shop floor could sit unresolved for most of a shift, and there was no tiering or ownership model to make sure the tickets that stopped work were handled before the ones that did not.
Network visibility was limited to what the team could see once something had already broken. There was no continuous view of server, network or storage health, so warning signs were only visible in hindsight.
The result was a function permanently in reaction. Every hour went to restoring what had failed, and none went into finding out why it kept failing
What the assessment revealed
Continuous monitoring changed what the team could see, and the pattern behind the failures became visible within weeks.
The failures were not unrelated. Incidents that had been logged and closed individually for more than a year shared a common source. Without monitoring data across the estate, there had been no way to see the connection.
Ageing storage hardware was reaching end of life. The storage supporting both production sites was the origin of the recurring server failures. It had been treated as a series of separate incidents rather than one asset approaching the end of its service life.
Alerting was buried in noise. Thresholds sat at vendor defaults rather than anything matched the client’s actual traffic patterns, so genuine warnings competed with routine alerts that required no action.
Patching depended on someone noticing. Updates were applied when a team member had capacity and remembered they were due, with no defined schedule and no record of what had been applied where.
Our approach
1. Monitoring and visibility, weeks one to three
We stood up 24/7 infrastructure monitoring across servers, network devices and storage, giving the team a single view of the estate for the first time.
Alert thresholds were tuned against the client’s actual traffic patterns rather than being left at vendor defaults, which cut the noise that had previously buried real warnings.
2. Helpdesk restructuring, weeks two to four
Support moved onto a tiered, SLA-based model with clear escalation paths and ownership at each tier.
Response times were tracked against the new SLA from week one, so performance against the commitment was visible from the point the model went live rather than assessed after the fact.
3. Root cause analysis, weeks three to five
With monitoring data available across both sites, recurring failures were traced back to ageing storage hardware nearing end of life.
This had been treated as a series of unrelated incidents for over a year. Once the pattern was established, the remediation became a single hardware decision rather than an open-ended troubleshooting exercise.
4. Hardware replacement and capacity planning, weeks five to eight
The storage hardware identified as the source of recurring failures was replaced, sequenced around production hours to avoid further disruption.
A capacity plan was put in place covering the following 18 months, with patch management brought under a defined schedule, so updates stopped depending on someone noticing they were overdue.
Technology stack
- Infrastructure monitoring: 24/7 monitoring across physical and virtual servers, network devices and storage, with thresholds tuned to the client’s own traffic baselines
- Service management: ITSM ticketing platform with tiered queues, defined escalation paths and SLA tracking and reporting
- Server and virtualization: Microsoft Windows Server, VMware vSphere, Microsoft Active Directory
- Storage: replacement shared storage across both production sites, sized against the 18-month capacity plan
- Patch and configuration management: scheduled patching with maintenance windows aligned to production hours, and configuration baselines recorded per site
- Supporting security work: endpoint protection coverage reviews and access control tidy-up delivered alongside the infrastructure programme
- Sites covered: two production sites, shop floor devices, back-office users and remote access
The results
65% reduction in unplanned downtime. Unplanned downtime fell by 65% within two quarters of the program starting, measured against the client’s own incident records for the preceding period.
Helpdesk response under 30 minutes. Average response time fell from four hours to under 30 minutes after the tiered SLA model was introduced, so shop floor faults stopped consuming most of a shift.
Zero recurring server incidents. No further recurring server incidents were recorded once the ageing storage hardware was replaced, ending a failure pattern that had run for more than a year.
100% SLA compliance. Every ticket handled after rollout was resolved within its committed SLA, with performance reported monthly rather than assessed subjectively.
Patching moved to a defined schedule. Updates now run against a fixed maintenance calendar aligned to production hours, with a record of what has been applied at each site.
Capacity planned 18 months ahead. The client moved from replacing hardware when it failed to knowing when it would need replacing, with a plan covering both sites.
All figures were measured from the client’s own incident and service desk records across the periods stated. Supporting details are available under NDA to qualified prospects.
In the client’s words
“We had been closing the same kind of ticket for a year without anyone joining the dots, because we never had a view of the whole estate at once. Once the monitoring was in and the alerts were tuned to us rather than to the vendor's defaults, the pattern was obvious within a fortnight. Replacing the storage was the easy part.”
IT Manager, mid-market manufacturer
Why it worked
Root cause analysis meant the ageing storage hardware was replaced once, rather than the symptoms being patched repeatedly. A year of individually logged incidents had cost more production time than the replacement did.
The sequencing mattered as much as the work itself. Monitoring came first because the root cause could not be established without it, and the hardware decision came last because it depended on the evidence produced during the first three weeks.
Tuning the monitoring thresholds to the client’s own traffic patterns meant the alerts that mattered stood out from day one. Default thresholds would have reproduced the same noise problem in a new tool.
The tiered helpdesk model gave every ticket to a named owner, which is what closed the gap between a four-hour average and a 30 minute one. The change was ownership, not effort.
Capacity planning and scheduled patching then moved the function out of reaction permanently. The team stopped depending on someone noticing that something was overdue.
Why 12th Wonder
Manufacturing environments make infrastructure availability a production issue rather than an IT one, which means stabilization work has to be judged on shop floor impact rather than ticket volume.
12th Wonder delivers IT Infrastructure and Support Services alongside Managed Security and Compliance, so monitoring, service desk, hardware lifecycle and security coverage are managed together rather than by separate parties with separate views of the same estate.
We hold no reseller or partner resale arrangements. The recommendation to replace the storage carried no commission to us, which is the only reason a hardware recommendation from a managed services provider is worth anything.
For this client, our team stayed involved from initial monitoring through root cause analysis, hardware replacement and into the ongoing managed service. The people running the environment today are the people who found the fault.
With more than 13 years of experience supporting enterprise IT environments, 12th Wonder provides managed services across cloud, infrastructure, security and IT operations.
We are ISO 9001 certified, NMSDC registered and have delivered more than 100 enterprise projects.
Services applied
IT Infrastructure and Support Services, with supporting work from Managed Security and Compliance.
Frequently asked questions
How long does an infrastructure stabilization program take?
This engagement ran eight weeks from initial monitoring deployment to completed hardware replacement and capacity planning, then transitioned into an ongoing managed service. The timeline depends on how much monitoring data exists at the start, since root cause analysis cannot begin without it.
Why start with monitoring rather than fixing the failures?
Because failures could not be diagnosed without them. Incidents had been logged and closed individually for over a year, and the common cause only became visible once there was a continuous view across servers, network and storage at both sites.
What does a tiered SLA-based helpdesk actually change?
It gives every ticket a named owner and a defined escalation path, so work that stops production is handled ahead of work that does not. In this engagement it took average response time from four hours to under 30 minutes without adding headcount.
Why tune alert thresholds instead of using vendor defaults?
Vendor defaults are set for a generic environment, not for a specific one. Left unchanged, they generate routine alerts that require no action, which is how genuine warnings get buried. Tuning against the client’s own traffic baselines is what makes an alert meaningful.
Can a small internal IT team sustain this after the project ends?
The monitoring, alert triage and service desk moved to 12th Wonder as an ongoing managed service, which is what allows a small internal team to stop being consumed by reactive work. Scheduled patching and an 18-month capacity plan removed the tasks that had previously depended on someone remembering them.
How is unplanned downtime measured?
Against the client’s own incident records for the period before the program, using the same definitions and the same reporting, so the before and after figures are comparable.
