Field notes · 15 min read ·
ShareManufacturing VMware exit: the plant floor sets the plan
Windows Server 2019 has been in extended support since January 2024 — security-only updates, ending January 9, 2029 — and post-Broadcom renewal pressure is arriving well ahead of that. On a plant floor the binding number is not the deadline. It is how many production windows sit between now and it.
If you run IT for a mid-market manufacturer, you have probably been handed a version of the same instruction twice in the last year: get off VMware, and get off Windows Server 2019. Both are reasonable. Both are, on paper, well-understood engineering. And both will run straight into a plant floor that does not care about either one.
This is the Field Notes read on what a plant floor actually does to a server modernization — the OT/IT split treated as a scheduling and risk constraint rather than an org chart. The split is not primarily about who reports to whom. It is about which systems you are allowed to stop, when you are allowed to stop them, and who has to sign for the consequences if the restart goes badly. Get that right and the migration is ordinary work. Get it wrong and you find out during a shift.
The cutover window is the constraint, not the technology
In an office estate, change is limited by engineering capacity. There is always another weekend. Scope grows, weekends get consumed, and the project finishes late but finishes. That intuition does not survive contact with a production line.
A line that cannot stop is a line that cannot stop. When a plant is running against booked orders, the maintenance window is not negotiated against user tolerance or a change-advisory calendar — it is negotiated against a production schedule that has already been sold. The windows that do exist are structural: an annual or semi-annual planned shutdown, a holiday period, a changeover between product runs, a seasonal demand trough, sometimes a gap created by a supply constraint upstream. They are set months in advance, they are contested by mechanical, electrical, safety and facilities work that also needs the line down, and IT is rarely first in that queue.
Two things follow, and they reshape the plan more than any platform choice.
First, you schedule backward. The question is not "how long will this migration take?" but "how many usable windows are there between today and the date I have been given, and what fits in each one?" If the answer is three windows and forty workloads with plant dependencies, that is your project plan, and no amount of parallelism at the hypervisor layer changes it. This is also why manufacturing VMware exits get quoted in months rather than weeks — the elapsed time is dominated by waiting for windows, not by consuming engineering hours. We describe the same shape on the manufacturing IT page: phased migration with cutover windows aligned to the production schedule, not to the project's convenience.
Second, the rollback has to fit inside the window too. A cutover that takes six hours of an eight-hour window, with a rollback that takes four, is not an eight-hour job. The genuine go/no-go decision point is at hour four, and if nobody has written that down before the night starts, the decision gets made at hour seven by whoever is most tired. Naming the decision point in advance — with the specific observable that triggers it — is the single cheapest risk control available on this kind of work.
There is a third effect that surprises IT teams the first time: the window is not the outage. Bringing a line back has its own ramp. Operator stations reconnect, interfaces and drivers re-handshake, buffered process data backfills, in-progress batch or work-order state has to be re-established, and quality usually wants a first-article check before the run is trusted. IT tends to declare success when the virtual machine responds. The plant declares success when the first good part comes off the line. Plan to the second definition.
What sits on the OT side that IT does not control
The uncomfortable part of a manufacturing hypervisor migration is that a large share of the operations-technology estate is already running on IT's cluster. It was virtualized years ago because that was the sensible thing to do, and it has been inside your VMware footprint ever since — which means it is inside your VMware exit whether or not anyone put it in the scope document.
What we consistently find on the far side of the line, described as categories rather than products:
- The process historian and its collectors. The time-series record of what the plant did. Continuously written, continuously read, and subject to retention obligations measured in years.
- SCADA front-ends and HMI servers. The systems operators actually look at. When these are down, the line is either stopped or running blind, and running blind is usually the worse of the two.
- Batch, recipe and MES tiers. Application and database servers that hold what the plant is supposed to be making right now, and the record of what it made.
- Engineering file services and program distribution. Drawings, machine programs, and the servers that feed programs to machine tools. Often the oldest file service in the building and the one with the most fragile permissions history.
- Machine-attached workstations. Physical, vendor-supplied, pinned to a specific operating system build, frequently connected by legacy serial or fieldbus interface cards, and sometimes dependent on a hardware licence key. These are not virtualization candidates and treating them as such wastes a window.
- Licence servers and entitlement bindings. Engineering, quality and controls software with licences bound to a hardware address, a machine identity, or a specific host. Migration changes those identifiers, which turns a routine move into a vendor ticket with its own lead time.
- Auxiliary services that stop production when they fail. Label and barcode printing, vision and test stands, gauging and data-capture nodes. Nobody lists them as critical until shipping cannot produce a label.
Ownership is the part that catches project managers. Many of these systems were bought on a capital equipment line, are supported by the equipment supplier rather than by IT, and are administered by controls engineering or plant maintenance. IT may hold the hypervisor and the network and still not hold the authority to reboot the guest. So the inventory that matters has three columns, not one: what the system does on the floor, who owns it commercially, and what stops if it is unavailable. Subnet membership tells you almost nothing.
The vendor-support trap: "just upgrade it" is a commercial decision
The single most common blocker on a manufacturing modernization is not technical difficulty. It is a support agreement.
Capital equipment is frequently sold with a support arrangement that names a validated software stack — an operating system build, driver versions, sometimes a specific virtualization platform and version. Moving outside that stack does not break the machine. It breaks the relationship. If the machine faults at 2am, the supplier's first question is what changed, and a modernized platform is an easy answer for them and an expensive one for you. In plants with formal validation or customer-audited processes, the cost is more concrete still: requalification is a documented exercise with protocols to re-run, first-article inspection to redo and change records to update, and that cost lands on quality and operations rather than on the IT budget.
So the decision is commercial. Three outcomes are available, and the work is choosing between them deliberately:
- Get a current supported configuration in writing. Often one exists and nobody has asked in five years; sometimes it exists behind a paid upgrade to a newer control package. Either way, the supplier's written supported-configuration statement is the document that most changes the plan, and obtaining it takes weeks. Start it during inventory, not during design. A project that reaches the design phase without those statements is designing against assumptions.
- Requalify deliberately. Move to a current platform and accept the validation work, with the plant's agreement and with the schedule cost booked where it actually falls. This is the right answer more often than IT expects, particularly when the same machine is due for a control upgrade anyway.
- Do not move it — isolate it. Leave the pinned host where it is and make it a documented, permanent exception: strict segmentation, no inbound path from the corporate network, brokered and recorded supplier access, removable-media control, and a rehearsed bare-metal restore so the box can be rebuilt when the hardware fails. This is a legitimate engineering outcome, not a failure of the project. What is not legitimate is leaving the machine broadly reachable, calling it accepted risk, and never writing down who accepted it.
The same logic applies to the hypervisor layer. When a supplier's supported stack names a virtualization platform, your destination choice for that workload is constrained before you open a feature comparison. That is one of several reasons we treat destination selection as a constraint-driven exercise rather than a product bake-off — the reasoning is laid out in the destination-platform decision framework, and the comparative tradeoffs across platforms are in VMware after Broadcom: exit options.
Air-gapped and semi-gapped segments
Modern infrastructure quietly assumes reachability. Activation talks to a licensing service. Patching pulls from an update channel. Backup agents check in. Management planes increasingly live in a cloud tenant. Certificate validation wants a revocation endpoint. Monitoring assumes an outbound path. On a genuinely isolated segment, every one of those assumptions needs an offline equivalent, and building those equivalents is the modernization — it is not preparation for it.
Concretely, a segment that cannot reach the internet needs: a local activation and licensing mechanism; a local update source plus a defined, controlled process for getting content into the segment; a backup target on the same side of the boundary and a rehearsed way to get a copy off it; either an in-segment monitoring collector or an explicitly accepted blind spot; and identity, name resolution and time services that keep working with the wide-area link down. That last one deserves emphasis. If a line cannot start when the corporate link is out, you have a production dependency hiding inside an IT design. Authentication and time belong physically inside the segment that has to keep running, and timestamps matter more here than in an office — historians, batch records and sequence-of-events logs are correlated on them during quality investigations.
The harder case, and the more common one, is the semi-gapped segment. It is documented as isolated and it is not. Over years it accumulates exceptions: a supplier's remote-support tunnel opened for a commissioning visit and never closed, an outbound proxy allowance for an update service, a dual-homed engineering workstation with a foot in both networks, a forgotten address-translation rule. The exceptions are the real architecture; the network drawing is a historical document. A modernization is the moment those exceptions surface, because you are about to change what sits on the other end of every one of them, and each will fail in its own way at its own time.
Enumerate them before you design, not during the cutover. In practice that means reading firewall rules and proxy allowances rather than reading the drawing, and asking every equipment supplier how they connect for support — the answer is frequently a path IT did not know existed.
Sequencing: the historian usually decides the plan
Once the constraints are known, sequencing is where the engineering judgement lives. The general principle is to move outward from the things nothing depends on toward the things the line stops for, proving the destination platform on low-consequence workloads before anything operational touches it.
A sequence that holds up in most plant estates:
- Non-production and office workloads. Not because they matter least, but because their purpose here is to prove the destination pattern — backup, restore, patching, monitoring, alerting — before an operational workload depends on it. The first plant system should land on a platform whose restore has already been executed, not described.
- Identity, name resolution, time and certificate services. Everything downstream authenticates, and the plant's records correlate on time. Placement matters as much as version: these services need to survive a link failure on the side of the boundary that has to keep producing.
- Operations-adjacent infrastructure that is not in the live data path. Licence servers, engineering file shares, access brokers, jump hosts. Disruptive if wrong, but they do not stop a running line the moment they blink.
- The historian, in a window of its own. See below — this is the one that sets the shape of everything around it.
- Systems the line genuinely stops for. SCADA front-ends, HMI servers, batch and recipe, MES. One system, one line, one plant at a time.
- Pinned and physically constrained machines, last. By this point you know what the rest of the estate looks like, which is exactly when an isolate-versus-requalify decision should be made rather than guessed.
Sequence plant by plant rather than by workload type across all plants. Migrating every historian across four sites in one window looks efficient on a Gantt chart and turns a bad night into a company-wide event. Keeping the blast radius to one plant keeps a rollback to one plant.
Why the historian decides. It is the system of record for process data — used for quality investigations, customer complaints, contractual or regulatory retention, performance reporting, and in some plants for batch release. It cannot simply be paused, because a gap in the record is the kind of thing somebody has to explain later. It is also usually the largest and most write-heavy dataset in the estate, so its migration method — replicate, copy and catch up, or rely on collector buffering — is the longest-lead engineering item in the project.
That gives you the number the whole plan hangs on. Most collection layers buffer locally when the historian is unreachable and backfill on reconnect. Your maximum tolerable cutover is approximately the smallest collector buffer, minus the time backfill itself takes. Buffers are usually sized in hours and usually left at whatever was configured at commissioning. So: pull the configured buffer depth for every collector, confirm it against the disk actually available on the collecting node, and prove it with a controlled disconnect before you plan around it. If nobody in the building knows the number, that is the first test to run — and finding it is often the most valuable single output of the discovery phase, because it silently caps the cutover length of everything upstream.
Where the plant estate quietly touches compliance
The manufacturing page names NIST SP 800-171, and it is worth being precise about when that is in scope, because it is frequently over-applied and occasionally missed entirely.
800-171 is scope only if you are in a defence supply chain. It applies to non-federal systems handling Controlled Unclassified Information, and it reaches manufacturers through DFARS flow-down from primes — including suppliers two or three tiers down who do not think of themselves as defence contractors at all. If your shop is purely commercial, with no defence work and no export-controlled technical data, it is not required, and building to it anyway is spending money on the wrong problem.
A migration does not create that scope. What it does is move the data that determines the boundary. Drawings, technical data packages and quality records get copied to new file services during a modernization. Backups land on a new target. A new management plane gains administrative reach into hosts that hold regulated data. Supplier remote-access paths get rebuilt. Each of those either sits inside a compliance boundary or outside it, and that is decided by choices made early — before the file services move, not after somebody notices during an audit that the backup repository is not where the paperwork says it is. If you are in scope, draw the boundary in the design phase and let it constrain the migration, rather than discovering it retroactively.
Two honest limits on our side of this. Pro IT NW performs readiness engineering for clients — we validate controls, document the boundary, and build the evidence a self-attestation rests on. We are not an assessor, and we hold no certification of our own. And whether a specific drawing or item is export-controlled is a trade-compliance determination belonging to your counsel or empowered official, not to your infrastructure consultant. We engineer the segregation boundary; we do not classify what goes inside it.
Commercial manufacturers are not exempt from all of this, incidentally — the same evidence gets demanded through customer security questionnaires flowed down by large buyers, and the underlying controls are largely the same. The difference is who is asking and what happens if the answer is weak.
What a workable plan looks like
Pulling the above together, the plan that survives contact with a plant floor tends to have the same seven elements:
- An inventory with an operations-adjacency flag, a commercial owner, and a "what stops" column. Built by walking the floor with controls engineering, not by exporting a list from the hypervisor.
- Written supported-configuration statements from every equipment supplier in scope. Requested in week one, because they take weeks and they determine which workloads are even movable.
- The collector buffer number, measured. It sets the maximum cutover length for everything writing into the historian.
- An eighteen-month shutdown and changeover calendar from operations. This is the project's actual capacity. Design against it, not against a start date.
- Destination selection driven by constraints. Device pass-through, licence binding, behaviour with the management plane unreachable, and supplier-supported platforms — before feature comparison. The VMware exit and server modernization service page covers how we run that assessment vendor-neutrally, without resale margin attached to the answer.
- A timed rollback with a named decision point. Written before the window, with the specific observable that triggers it.
- A definition of done that says "first good part". Not "guest responds to ping".
Two more sequencing notes worth stating plainly. If Server 2019 or Server 2016 upgrades are also on your list, do them coincident with the hypervisor move wherever the workload allows — you are paying for the window either way, and a second stop for the same server is a second negotiation with the plant. The options and the tradeoffs are in Windows Server 2019 end of support: upgrade options and Windows Server 2016: extended security updates or modernize. And resist the temptation to solve the pinned machines first because they are the most interesting problem. They are the last decision, not the first, because their answer depends on what the rest of the estate ends up looking like.
Related reading
- The industry view, including where we engage and where we deliberately stop: Manufacturing IT — VMware exit, Server 2019 EOL, 800-171.
- The service pillar and how destination selection is run without resale margin: VMware Exit & Server Modernization.
- The structured way to choose a destination platform: Destination-platform decision framework.
- Comparative platform tradeoffs after Broadcom: VMware after Broadcom — exit options.
- Server 2019 upgrade paths: Windows Server 2019 end of support — upgrade options.
- Server 2016 and the extended-security-updates question: Windows Server 2016 — ESU or modernize.
The 30-second version
On a plant-floor estate, the hard constraint on a VMware exit or a Server 2019 upgrade is not the technology — it is the production schedule. Count the usable windows between today and your deadline; that is the project's real capacity, and the rollback has to fit inside each one. A large share of the operations-technology estate is already virtualized on your cluster and is therefore in scope whether or not it was written down, but IT frequently does not own it commercially. Equipment support agreements that pin an operating system or a hypervisor version make "just upgrade it" a commercial negotiation, with three legitimate outcomes: get a current supported configuration in writing, requalify deliberately, or isolate the machine permanently and document who accepted that. Isolated and semi-isolated segments need offline equivalents for activation, patching, backup, monitoring, identity and time — and the semi-gapped case is the dangerous one, because its accumulated exceptions are the real architecture. Sequence outward from workloads nothing depends on to the ones the line stops for, plant by plant, and let the historian's collector buffer set your maximum cutover length. NIST 800-171 is scope only if you sit in a defence supply chain — but if you do, decide the data boundary before the file services move.
If you want a senior engineer to walk your plant estate and turn it into a windowed migration plan, the project intake form takes about three minutes. Tell us the plant footprint, the workloads, and the deadline you have been given, and we will come back with a destination recommendation and a windowed migration scope. More detail on how we work with manufacturers is on the manufacturing IT page.
If you want a senior engineer to scope a VMware exit or a Server 2019 upgrade against your actual plant-floor constraints, the project intake form takes about three minutes. Two-business-day response with scope and a fixed-fee range.
Pro IT NW engineers VMware exits, server modernization and Microsoft 365 work for mid-market discrete and process manufacturers. Senior-led, labor-only, vendor-neutral: we do not resell hardware, hypervisors or licensing, and we do not reconfigure PLCs, SCADA or MES — those belong to your operations technology team and your equipment suppliers. We engineer the platform and the boundary that meets them.
Questions we get asked
- Why is a VMware exit harder at a manufacturer than at an office-only business?
- Because the scarce resource changes. In an office estate the constraint is engineering effort, and a weekend is usually available for any given workload. In a manufacturing estate the constraint is the production schedule: a line that is booked against customer orders has no maintenance window, and the windows that do exist — planned shutdowns, changeovers between product runs, holiday or seasonal troughs — are calendar events set months ahead and contested by mechanical, electrical, safety and facilities work at the same time. That inverts the plan. You do not schedule forward from a project kickoff; you schedule backward from a small, fixed number of windows, and the number of windows between today and your deadline is the real capacity of the project. The hypervisor migration itself is the easy part.
- Which plant-floor systems end up inside an IT server modernization even when nobody scoped them?
- More than most inventories show. Virtualized operations technology usually lives on the same cluster as the corporate workloads, so it is inside a hypervisor migration whether or not it was written into the scope. The recurring list: the process historian and its collectors, SCADA front-end and HMI servers, batch and recipe servers, MES application and database tiers, engineering file shares holding drawings and machine programs, DNC or program-distribution servers feeding machine tools, label and barcode print servers, vision and test-stand controllers, and licence servers whose entitlements are bound to a hardware address or a virtual machine identity. Then there are the machine-attached workstations that are not virtual at all and never will be. The practical rule is to inventory by what the system does on the floor and who owns it, not by which subnet it happens to sit on — ownership frequently sits with controls engineering, plant maintenance or the equipment supplier rather than with IT.
- Our equipment vendor says changing the operating system voids support. What are the real options?
- There are three, and choosing between them is a commercial decision rather than a technical one. First, get the vendor to state a supported configuration that includes a current platform — sometimes one exists and nobody asked, sometimes it exists behind a paid upgrade to a newer control package. Second, requalify: move to a current platform, accept the validation or first-article work that follows, and absorb the schedule cost that lands on quality and operations rather than on IT. Third, do not move it — isolate it instead, and treat the pinned host as a permanent, documented exception with compensating controls: strict segmentation, no inbound path from the corporate network, brokered and recorded remote access, removable-media control, and a tested bare-metal restore so the box can be rebuilt when the hardware eventually fails. Isolation is a legitimate engineering outcome, not a failure. What is not legitimate is leaving the machine reachable and calling it accepted risk without writing down who accepted it.
- How long an outage can a process historian actually tolerate?
- That is a number to extract from the environment rather than estimate, and it is usually the single most useful figure in the whole plan. Most collection layers buffer locally when the historian is unreachable and backfill on reconnect, so the tolerable outage is roughly the smallest store-and-forward buffer among the collectors, minus the time the backfill itself takes. Buffers are commonly sized in hours, and they are commonly sized to a default nobody has revisited since commissioning. Ask for the configured buffer depth per collector, confirm it against the disk actually available on the collecting node, and test a controlled disconnect before you plan around the answer. If nobody in the building knows the number, finding it is the first test of the project — because it silently caps the cutover length of everything that writes into the historian.
- Can you modernize a segment that genuinely has no internet access?
- Yes, but every convenience that assumes a reachable service has to be replaced with an offline equivalent, and that work is the project rather than an afterthought. Activation and licensing need a local mechanism. Patching needs a local update source with a defined path for getting content into the segment. Backup needs a target on the same side of the boundary, plus a rehearsed way to get a copy off it. Monitoring and management need either an in-segment collector or an accepted blind spot. Identity and time need to work with the wide-area link down, which usually means authentication and time sources physically inside the segment — if a line cannot start when the corporate link is out, the design has a production dependency hiding in it. The more common and more dangerous case is the semi-gapped segment: nominally isolated, but carrying years of accumulated exceptions such as supplier remote-support tunnels, outbound proxy allowances and dual-homed workstations. Enumerating those exceptions is part of the modernization, because you are about to change what sits on the other end of them.
- What order should systems move in during a plant-floor server modernization?
- Move outward from the things nothing depends on, and end at the things the line stops for. A defensible sequence: non-production and office workloads first, purely to prove the destination platform and its backup, restore, patching and monitoring pattern; then identity, name resolution, time and certificate services, because everything downstream authenticates and because historians, batch records and sequence-of-events logs correlate on timestamps; then operations-adjacent infrastructure that is not in the live data path, such as licence servers, engineering file shares and access brokers; then the historian, in a window of its own sized against the collector buffer; then the systems the line genuinely stops for — SCADA front-ends, HMI servers, batch and recipe, MES — one system, one line and one plant at a time; and last the pinned or physically constrained machines, whose answer is usually architectural isolation rather than migration. Sequence plant by plant rather than workload type across all plants, so that a bad night stays a single-plant problem.
- Does a VMware exit or server upgrade put a manufacturer in scope for NIST 800-171?
- Only if the business is already in scope, and a migration does not create scope by itself. NIST SP 800-171 applies to non-federal systems handling Controlled Unclassified Information, which reaches manufacturers through DFARS flow-down from defence primes — including suppliers several tiers down who do not think of themselves as defence contractors. A purely commercial manufacturer with no defence or export-controlled work is not required to meet it. What a migration does change is where regulated data lives: engineering drawings, technical data packages and quality records get copied to new file services, new backup targets and a new management plane during a modernization, and each of those can land inside or outside a compliance boundary depending on decisions made early. If the business is in scope, define the boundary before the file services move, not after. Pro IT NW performs readiness engineering for clients; we are not an assessor and we hold no certification of our own.
- Do you reconfigure PLCs, SCADA or MES as part of this work?
- No. Programmable controllers, SCADA configuration, historian tag structure and MES application logic belong to the operations technology team or the equipment supplier, and Pro IT NW does not change them. What we engineer is the boundary and the platform the boundary rests on: the virtualization and server estate those systems run on, identity and access for operations-adjacent accounts, segmentation where the plant network meets the corporate network, brokered remote access for suppliers, file services for drawings and quality records, and the backup and restore design. Boundary discipline is deliberate — in manufacturing, consulting scope creep past that line is how a project damages production.
- What counts as a successful cutover on a plant floor?
- The first good part off the line, not a server responding to a ping. Restart has a ramp that IT change records rarely capture: operator stations reconnect, drivers and interfaces re-handshake, buffered data backfills, in-progress batch or work-order state has to be re-established, and quality typically wants a first-article check before the run is trusted. Plan that ramp into the window as explicitly as the migration steps, and set the go or no-go decision point early enough that a full rollback still fits inside the same window. A cutover that consumes six hours of an eight-hour window with a four-hour rollback is not an eight-hour job — the real decision point sits at hour four, and naming it in advance is what keeps an overrun from becoming an unplanned production stop.
Related service
Manufacturing ITWritten by the team at Pro IT NW · Senior-led Microsoft project consultancy · Seattle / USA-wide.