TLDR: Muse Glimmer, 30b, dense, 24 GB VRAM.
Meta's first track all the way back from 2023 is a huge part of the open source's community.
We've got llama.cpp, ollama and so many more. Safe to say there were one of the firsts and at the time, the best at it.
It is a 30b model, dense so it won't be the fastest. Good news is that they've optimized it for "consumer GPU's, aka a 3090 at least (24 GB VRAM). The 4 bit quantization runs on about 20 GB's of VRAM with 4 GB remaining for the KV cache, aka longer convos.
The bench marks do speak for themselves, they claim it's the best model in it's class i.e. 30b params or so, but testing it for your use case is much more safer.