To enable GPU acceleration, you must explicitly instruct llama.cpp to offload model layers to your graphics hardware.
-ngl (number of GPU layers) flag to your command. For example, to offload all layers, use: -ngl 999.nvidia-smi or an equivalent GPU monitoring tool during inference.
Comments
Sign in to join the conversation
Sign InNo comments yet. Be the first to share your thoughts!