Inheriting an Undocumented IT Estate: Where to Start
Taking over an estate nobody documented is a common situation and a solvable one. The mistake is trying to document everything at once, which never finishes. Working in the right order gets you safe quickly and complete eventually.
Start with what an incident would need
Not a full inventory. The question is: if something failed tonight, what would you wish you had written down?
Four things, in order.
Photographs of every rack, front and rear. Faster than notes and they capture things nobody would think to record. Date them.
Out-of-band access. Can you reach management controllers and console ports, and on a path independent of the production network? If not, that is the first gap to close β our out-of-band guide covers what makes it genuinely independent.
Credentials. Management controllers, switches, storage, UPS. Stored somewhere that survives the loss of the systems they unlock, which our documentation guide covers.
What fails if this goes down. Even a rough map of dependencies is worth more than a precise inventory of things nobody uses.
That is a day or two of work and it changes your position considerably.
Let the machines tell you what they are
The estate can describe itself, and most people do this the slow way.
Management controllers report model, serial, installed processors, memory, adapters, drive status, power supply state and firmware levels β without opening anything or taking anything down. Our management guide covers access.
Storage controllers report array configuration and drive health, including part numbers and power-on hours β our drive health guide covers reading them.
Managed switches report port status, negotiated speeds, error counters and what is connected where.
Collecting that is far quicker than a physical audit and gives you most of what matters. Do the physical walk afterwards, to confirm and to find what is not reporting.
Find the things nobody mentioned
Every inherited estate has them.
Machines with no owner. Running, consuming power, and nobody knows why. The technique from our decommissioning guide applies: look at what connects to what, and where practical make something temporarily unavailable and see who complains.
Scheduled jobs. The monthly or quarterly process that will fail long after you concluded the audit was finished. These do not show up in a week of observation.
Equipment in odd places. A switch in a cupboard, a small server under a desk, a device in a ceiling void. Our branch cabinets guide covers why those locations fail β heat, usually.
Isolated systems deliberately not on the network, which by definition will not appear in any network-based discovery. Ask, rather than scan β our isolated networks guide covers why they exist.
Check the things that fail quietly
An inherited estate has been running without anyone watching for warnings, so some are already showing.
Drive health attributes across every array. Rising reallocated or pending counts mean a replacement is due, and correlated ageing means several drives may show it at once.
Power supply status. A server running on one supply because the other failed months ago is common and is a redundancy failure nobody noticed β our power supply guide covers why that is urgent.
Fan status and inlet temperatures. A machine already throttling looks like a slow machine rather than a fault, as our fans guide covers.
Cache battery status on RAID controllers β a failed one means write-through mode and poor performance with no error, which our controller guide covers.
UPS runtime. Batteries age whether used or not, and an inherited UPS is frequently protecting nothing β our UPS battery guide covers why self-tests mislead.
Backups. Not whether jobs report success, but whether a restore actually works. Our backup guide covers why an untested backup is an assumption.
That list frequently surfaces two or three things needing attention in the first week.
Record what constrains each item
The step that turns an inventory into something useful.
For each significant item: what it is and its part number, roughly when it entered service, what it runs, what depends on it, and what would force its replacement.
That last field is the one that matters, and our lifecycle guide covers why. On an inherited estate the answer is frequently "software support ended two versions ago" or "parts are getting hard to find" β both worth knowing before they become urgent.
Part numbers matter most on older platforms. Our part numbers guide covers reading them from fitted components, and it is worth doing while machines are healthy rather than during an outage.
What to buy first
Once you know what you have, spares are the fastest improvement to your position.
Weight them by consequence rather than evenly: drives matching what is in production arrays, a power supply per chassis type, fans for machines that shut down without them, and a boot device where one is used. Our spares guide covers building the holding.
On platforms where parts are thinning, buy sooner rather than later β our availability guide covers why waiting makes that worse rather than better.
Keep it current from here
Two habits stop you inheriting the same problem again.
Update at the point of change rather than afterwards. A change that includes updating documentation as a step gets done.
Label both ends of any cable you touch. Within a year most of the estate is labelled without a project, which is the approach our documentation guide recommends.
And store all of it somewhere that survives the loss of the infrastructure it describes. Documentation on the server that is down is unavailable exactly when it is needed.
Common questions
Where do I start with an undocumented estate?
With what an incident would need rather than a full inventory: rack photographs, working out-of-band access, credentials stored where they survive the systems they unlock, and a rough dependency map. That is a day or two and it changes your position.
Do I need to physically audit everything?
Not first. Management controllers, storage controllers and managed switches report model, serial, components, health and firmware without opening anything. Collect that, then do the physical walk to confirm and to find what is not reporting.
How do I find machines nobody knows about?
Look at what connects to what, and where practical make something temporarily unavailable and see who complains. Watch particularly for scheduled jobs that run monthly or quarterly β those will not appear in a week of observation.
What is usually already wrong?
An estate that ran without anyone watching for warnings usually has some showing β a server on one power supply, drives with rising error counts, a failed cache battery causing slow writes, a UPS with aged batteries, and backups nobody has restored from.
What should I record about each item?
Beyond what it is: its part number, roughly when it entered service, what depends on it, and what would force its replacement. On an inherited estate that last answer is frequently ended software support or thinning parts availability.
Send us the part numbers you find on older platforms and we will tell you honestly how thin the market is for them.
