Who would benefit from it?
Straker operates multiple few-billion-parameter custom models and is evaluating IBM Cloud IKS as a hosting platform for these workloads. Straker would directly benefit from GPU-sharing capabilities that improve utilization and reduce the number of physical GPUs required.
This capability would also benefit partners and enterprise customers running multiple models that do not require an entire physical GPU per workload, enabling them to improve GPU utilization and control infrastructure costs.
Why is it useful?
On IBM Cloud IKS VPC GPU worker pools, the NVIDIA driver stack and GPU operator are platform-managed and not partner-configurable. Neither GPU time-slicing nor Multi-Instance GPU (MIG) partitioning can be enabled, so each model pod requires a separate physical GPU regardless of model size.
Straker currently serves multiple few-billion-parameter models on a single L40S-class GPU using time-slicing on another public cloud's managed Kubernetes platform. On IKS, the same models would each require a dedicated GPU.
For Straker's fleet of small models, this would increase the required GPU count by an estimated two to four times, leaving visible GPU capacity idle and unusable. This significantly increases infrastructure costs and makes IKS less economically viable for these workloads.
How should it work?
IBM Cloud IKS should support:
Alternatively, IBM Cloud should provide partner-configurable access to the NVIDIA GPU operator on dedicated GPU worker pools.
These capabilities should be supported and documented without requiring customers to make unsupported changes to the platform-managed NVIDIA driver stack.
| Idea priority | High |
| Needed By | Week |
By clicking the "Post Comment" or "Submit Idea" button, you are agreeing to the IBM Ideas Portal Terms of Use.
Do not place IBM confidential, company confidential, or personal information into any field.