There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
Like in the existing SIMD technology?
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_la...
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
- how exactly does apple silicon s unified gpu + ram thingy work?
- how come intel and nvidia cannot do the same?
- is there an actual difference in terms of hardware architecture or something or is it pure apple marketing hype?
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data. You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.
Why not make a faster interconnect with an open communication standard?
closed