Why This Has Been Expensive Until Now

VDI — Virtual Desktop Infrastructure — delivers a full desktop from a data centre or cloud instead of from the device in front of you. A user signs in from a laptop, a thin client or their own machine, and gets a familiar Windows or Linux desktop whose apps, files and processing all happen centrally.

The appeal for an IT team is that patching, access control and data protection all happen in one place, and people can reach the same desktop from anywhere.

The obstacle has never been the hypervisor. It has been the GPU.

Sharing one physical GPU between several desktops has been NVIDIA’s territory, and NVIDIA charges for the vGPU software that does the slicing. That licence is why VDI with hardware-accelerated graphics has mostly been the preserve of outfits with corporate budgets. Paying a yearly fee to switch on something the silicon can already do is a hard thing to be keen about.

Intel changed the arithmetic with the Arc Pro line. These cards support hardware splitting through standard SR-IOV, which is a PCIe feature rather than a product. As such, there is no licence server, no subscription, and no separate vGPU software to buy.

So I got hold of an Intel Arc Pro B50 to see how cleanly it goes together with Proxmox. The short answer: the mechanism works exactly as advertised, and then firmware got in the way. Both halves are below.

How one card becomes several: physical function on the host, virtual functions to guestsIntel Arc Pro Bxx03:00.0 — physical functionhost driver: xe03:00.1vfio-pci03:00.2vfio-pci03:00.3created where supported03:00.4created where supportedWindows desktop 1hardware acceleratedWindows desktop 2hardware acceleratedWindows desktop 3hardware acceleratedWindows desktop 4hardware acceleratedNo vGPU licence is involved at any point. SR-IOV is a PCIe capability the card either advertisesor does not — this one advertises it at [320]. How many functions it will create is set infirmware, because each one is given a fixed slice of the card's memory: a bigger slice meansfewer functions. A 16 GB B50 gives two at 8 GB each; the 24 and 32 GB cards report seven.
The physical function stays with the host’s xe driver. Each virtual function is a PCIe device in its own right, bound to vfio-pci and handed to a guest — the same passthrough machinery as a whole card, just several times over. The dashed pair depends on the card: how many functions each one will create is further down.

The Elephant: You Need Windows First

Before any of this works, the card wants its firmware updating. Intel ships that update inside the Windows driver installer.

Which is awkward, because the reason you bought the card is to run it under Proxmox.

There is a way round it that needs nothing but Proxmox: build a Windows VM, pass the whole card through to it, let Windows update the firmware, then hand the card back to the host. It is a bootstrap loop, and it is worth knowing about before you plan the build rather than after.

Breaking the Windows-first bootstrap loop with a temporary VMThe catch: SR-IOV needs current firmware, and the firmware ships inside a Windows driver installer.1Card in the hostlspci → 03:00.02Whole card → Windows VMhostpci0, Q35 + OVMF3Install Intel driverfirmware updates with it4Reboot hostcard reinitialisescard handed back to Proxmox5SR-IOV presentlspci -v → [320]6tmpfiles.d at bootnumvfs, unbind, bind7VFs to guestsas many as firmware allowsSteps 2 and 3 exist only to update firmware. The Windows VM is temporary — once the card is backon the host it plays no further part, and nothing about the finished setup depends on Windowsrunning on the hypervisor.
The card cannot do SR-IOV until its firmware is current, and the firmware update arrives as a Windows driver. One temporary Windows VM with the whole card attached breaks the loop.

Finding the Card

On the Proxmox host, lspci to locate it:

lspci

lspci output on the Proxmox host showing 03:00.0 VGA compatible controller: Intel Corporation Battlemage G21

Here it is at 03:00.0, reported as Battlemage G21, the silicon behind the Arc Pro B50. Note the address down. You will need it several times, and there is a separate audio function at 04:00.0 that comes with it.

Passing the Whole Card to a Windows VM

Add the card to a Windows VM as a raw PCI device. In the VM’s hardware, that is a PCI Device entry: 0000:03:00 with pcie=1:

Proxmox VM hardware tab showing 16 GiB memory, 4 host cores, OVMF UEFI, machine pc-q35-10.1, VirtIO SCSI single, a TPM state device, and PCI Device hostpci0 set to 0000:03:00 with pcie=1

Worth noticing what else that VM is, because none of it is accidental: Q35 machine type, OVMF firmware, VirtIO SCSI single, and a TPM for Windows 11. Q35 in particular is not optional for this. A passed-through GPU on i440fx appears as a legacy PCI device, which is the wrong shape for a modern graphics driver. That is covered in Always Use Q35, Not i440fx.

Boot the VM and check Windows sees the card:

Windows Computer Management, Device Manager, Display adaptors showing two Microsoft Basic Display Adapter entries, one flagged with a warning

It appears as a Microsoft Basic Display Adapter because no driver is installed yet. That is the expected state, and it is enough. Windows has found the hardware.

Updating the Firmware

Get the current driver from Intel’s Arc Pro B50 download page. At the time of writing that was version 32.0.101.8306 (Q4.25), for Windows 11 and Windows 10 22H2:

Intel’s Arc Pro B50 Graphics downloads page listing Intel Arc Pro Graphics for Windows, version 32.0.101.8306 Q4.25, dated 23 December 2025

Install it, and let the firmware update run as part of the process rather than cancelling out early. That firmware step is the entire reason for this detour.

When it finishes, give the card back to the host and reboot, so it reinitialises fully under Proxmox.

Confirming SR-IOV Is There

Now ask the card what it can do:

lspci -v

lspci -v output for 03:00.0 showing kernel driver in use xe, IOMMU group 13, and capabilities including Alternative Routing-ID Interpretation, Address Translation Service, Physical Resizable BAR, Virtual Resizable BAR and Single Root I/O Virtualization at 320

The line that matters:

Capabilities: [320] Single Root I/O Virtualization (SR-IOV)

That is the whole proposition in one line of lspci output, with no licence attached to it.

Three other things in that output are worth reading while you are there:

  • Kernel driver in use: xe — the card is on Intel’s newer xe driver rather than i915, which is what will be creating the virtual functions.
  • [420] Physical Resizable BAR and [220] Virtual Resizable BAR — the card supports resizable BARs, and so do its virtual functions. Worth knowing what that costs you in address space if you are passing several of them through: see PCIe Resizable BAR and Modern GPUs.
  • IOMMU group 13 — the card is in a group of its own, which is what you want for clean passthrough. Why that matters is in the IOMMU tax article.

Creating the Virtual Functions at Boot

Virtual functions are not persistent. Asking for them is a write to sysfs, so it has to happen on every boot.

tmpfiles.d is a tidy way to do that declaratively, rather than bolting a script onto a unit file:

cat /etc/tmpfiles.d/b50-setup.conf showing a write to sriov_numvfs, four unbind writes to the xe driver, and four bind writes to vfio-pci

The file does three jobs in order.

Create the virtual functions, by writing the count to the physical function’s sriov_numvfs:

w /sys/devices/pci0000:00/0000:00:01.1/0000:01:00.0/0000:02:01.0/0000:03:00.0/sriov_numvfs  - - - -  4

Unbind each new function from xe, because the host driver claims them as they appear and a guest cannot have a device the host is holding:

w /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.1
w /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.2
w /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.3
w /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.4

Bind them to vfio-pci, which is what makes them available to pass through:

w /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.1
w /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.2
w /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.3
w /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.4

Note the addresses: the physical function is 03:00.0 and the virtual functions come up as .1 through .4.

One detail about w that explains the ordering, and that will bite you if you get it wrong. systemd documents it as: “Write the argument parameter to a file, if the file exists.” sriov_numvfs only exists once a driver has bound to the physical function, and the virtual function paths only exist once that write has happened. So the sequence in the file is not stylistic. Each line depends on the one before it having taken effect.

Why Two, and Not Four

Driver 32.0.101.8306 — the one installed above — carries graphics firmware BMG__21,1162, and that is the release where Intel first officially enabled SR-IOV on Arc Pro. Intel’s stated default for the B50 in that release is two virtual functions, each with an 8 GB VF Local Memory BAR.

Which makes the number arithmetic rather than policy. The B50 has 16 GB. At 8 GB per virtual function, two is all that fits.

There is a wrinkle worth knowing if you go looking for a workaround. Before official support existed, some people ran older firmware that exposed 12 virtual functions on a B50, and going back to driver 32.0.101.6979 restores that count. Those 12 shared the same 16 GB, so each got a fraction of the memory mine get. Intel’s position is that two was chosen deliberately to give each function enough compute, capacity and bandwidth to behave predictably.

So the cap can be moved, but not by you. The maximum VF count and the VF Local Memory BAR size live in the IFWI, there is no public tool to change either, and the supported answer on a current stack is two.

How Many Desktops Each Card Gives You

The B50 is the small card in the family, and its two functions are the family’s low water mark. If the seat count is what you care about, buy further up the range.

The whole Battlemage Arc Pro line does SR-IOV. What differs is how many functions the firmware will carve out, and that follows the memory:

CardMemoryVFs on the current supported stackSeen elsewhere
Arc Pro B5016 GB2, at 8 GB VF BAR each — Intel’s documented default12 on pre-official firmware, via driver 32.0.101.6979
Arc Pro B6024 GB7 reported24 on an early ASRock firmware, cut to 7 by a later one
Arc Pro B60 Dual2 × 24 GB7 per GPU — two GPUs, so 14 from one slotas above; the two halves are independent
Arc Pro B6532 GBno published count found—
Arc Pro B7032 GB7 reported, on firmware 85174 on earlier firmware

Only the B50 row is Intel-documented. The B60 and B70 numbers are what people report from lspci, and they have moved more than once. The B60 in particular went from 24 down to 7 in a firmware update, which is the same kind of narrowing the B50 saw. Nobody appears to have published a VF count for the B65 at all, so treat that row as unknown rather than as zero.

Two things follow from that table, and both matter more than any single number in it.

The VF count is a memory division, not a die feature. Intel’s rule is that a larger VF Local Memory BAR means fewer functions. That is why the 16 GB card gives two and the 32 GB cards give seven: nothing about the GPU’s shaders decides it.

Check the card you are about to buy, not the family. SR-IOV presence has varied between board vendors on the same chip — Sparkle’s B60 Blower initially shipped without the capability visible at all and only gained it after an igsc firmware update. Ask for lspci -v output from the exact model, or budget for a firmware update before you count on any of this.

The Dual B60 Is Two Cards Wearing One Bracket

Maxsun’s Arc Pro B60 Dual 48G Turbo is the interesting one for seat count, and the thing to understand is that the 48 GB is not a pool.

It is two B60 GPUs — two BMG-G21 dies — on one board, with 24 GB of GDDR6 wired to each, and no PCIe bridge chip between them. Both dies hang straight off the x16 gold fingers at PCIe 5.0 x8 each.

Which means the host has to split the slot for you. The card needs the primary x16 slot bifurcated to x8/x8, and most consumer boards do not enable that by default. It is a firmware setting you go looking for, in the same category as the IOMMU and ACS settings any of this needs.

Get that right and the operating system sees two separate GPUs, each with its own physical function and its own SR-IOV capability. So you get two lots of virtual functions from one slot — 14 seats if each die behaves like a single B60 — and the tmpfiles.d file above doubles up, one sriov_numvfs write per die.

Get it wrong and you see one GPU and half the card is invisible.

Worth being plain about what the 48 GB is not: a guest attached to a virtual function on the first die cannot reach the second die’s memory. This is two 24 GB cards in one physical space, which is exactly what you want for VDI seats and exactly what you do not want for one large model.

So: if two seats are enough, the B50 is a 70 W card that will do it. If you want seven, plan around a B70. If you want fourteen and have a board that will bifurcate, the dual B60 gets you there in one slot.

Can You Run AI On a Virtual Function?

Short answer: treat it as unsupported. Longer answer, because the reason matters and it is not the one you would guess.

These are marketed as AI cards and they are not pretending. The B50’s 128 XMX engines are rated at 170 peak TOPS, the B70 at 367, and Intel’s software story is real — vLLM serves models from 8B up to 120B on Arc Pro B-series, and IPEX-LLM and llama.cpp’s SYCL backend both run on them.

But look at how every one of those results is produced. vLLM’s own Arc Pro numbers come from a Docker container on bare metal, on systems with four and eight whole B60 cards doing tensor parallelism. Intel’s post does not mention SR-IOV or virtual functions once.

That pattern holds everywhere I looked. Intel scopes the SR-IOV use cases to virtualised remote desktop, guest OS graphics acceleration, and media encode and decode. Compute is not on that list, and I could not find a single published case of anyone running LLM inference inside a VM attached to a virtual function.

What people actually do is telling: they run the model in Docker on the host, and hand virtual functions to VMs for desktops. One person doing both at once reports simply that “VRAM gets pretty tight.”

Which is the real problem, and it is arithmetic rather than driver support.

A virtual function gets a fixed slice of local memory — 8 GB on the B50, set in firmware. That slice is the hard ceiling for weights plus KV cache in that guest, and it does not grow because the card has more. An 8B model at FP16 is around 16 GB of weights before you add any context at all, so it does not fit in a B50 virtual function on any driver. Quantise to Q4 and an 8B fits in about 4 GB, leaving a few GB for context — which works, but is a long way from what the card can do undivided.

So the two workloads compete for the same memory, and the split is decided in firmware before either of them starts.

If AI is the job, do not divide the card. Pass the whole thing through to one VM — the same hostpci0 passthrough used for the firmware update earlier in this post — or run the container on the host and skip virtualisation for that workload. Both give the model all 16 GB and the full XMX array.

If VDI is the job, virtual functions are right, and expect desktop graphics rather than an inference server behind each one. Hardware-accelerated desktops, video playback and encode work. That is what the mechanism is documented for.

Worth saying plainly: absence of published evidence is not proof it fails. The xe driver exposes compute through Level Zero and OpenCL, and it is entirely possible a VF-backed guest brings those up fine. But nothing from Intel says it is validated, no one appears to have shown it working, and the memory ceiling limits the payoff even if it does. That is not something to build a plan on.

What It Is Still Worth

Two virtual functions is two hardware-accelerated Windows desktops from one card, with no vGPU licence, no subscription and no licence server. On a hypervisor that costs nothing to run. That is enough to prove the approach works, which is the honest job for a B50. It is the bottom of the range.

For an actual VDI deployment I would be specifying the dual B60.

Fourteen functions from one slot — if each die behaves like a single B60 — puts it in the same seat-count territory as the NVIDIA cards sold for this workload, at a far lower price, and with nothing to license per user. That last part is the one that compounds. NVIDIA’s vApps, vPC and RTX vWS are all licensed per concurrent user, either as an annual subscription or as a perpetual licence that has to be bought alongside a five-year support and maintenance subscription. Every seat is a line item, and it comes round again. On the Intel side there is no equivalent line. You buy the card.

And the mechanism is the part that matters long term. SR-IOV on the GPU is a PCIe capability, not a product tier, so the tmpfiles.d file just grows to match whatever the card allows.

The shape is one sriov_numvfs write, then an unbind and a bind for each function — so two functions is five lines, and the nine above are four functions asked for on a card that delivers two. Seven functions is fifteen lines. A dual B60 is thirty, because each die is its own physical function and gets its own sriov_numvfs write.

Nowt else changes as you scale it. No licence server appears at any point in that file.

References