Well, thing is, Nvidia doesn’t have to ask. They don’t need permission to ship it, and once it’s there anyway this system has a much higher chance of being adopted.
This is basically what they’ve done with past standards, like certain compute functions. They ship it, and then people end up using it.
A simple example is low precision data formats. The machine learning world went from largely working with 32 bit floating point to certain specifications of FP16, then BF16/tensor int8, then FP8, then FP4.
Why?
Because that’s what Nvidia shipped! Alternatives were proposed and entertained, yet they faded away because they didn’t have native hardware support. To be specific, FP16 was Pascal (GTX 1000 cards), int8 was sort of Turing/Volta (RTX 2000), BF16 was Ampere (RTX 3000), FP8 was Ada (RTX 4000), FP4 was Blackwell (RTX 5000).
Well, thing is, Nvidia doesn’t have to ask. They don’t need permission to ship it, and once it’s there anyway this system has a much higher chance of being adopted.
This is basically what they’ve done with past standards, like certain compute functions. They ship it, and then people end up using it.
A simple example is low precision data formats. The machine learning world went from largely working with 32 bit floating point to certain specifications of FP16, then BF16/tensor int8, then FP8, then FP4.
Why?
Because that’s what Nvidia shipped! Alternatives were proposed and entertained, yet they faded away because they didn’t have native hardware support. To be specific, FP16 was Pascal (GTX 1000 cards), int8 was sort of Turing/Volta (RTX 2000), BF16 was Ampere (RTX 3000), FP8 was Ada (RTX 4000), FP4 was Blackwell (RTX 5000).