Yes I do! 5060 is tough because you have 8GB of VRAM. That means any model you load onto it will have to be smaller than that (a model that is 6GB on hugging face will take up 6 GB of VRAM just loaded onto your GPU, then need some more headroom for the actual inference).
So, you could get away with Qwen3.8 quantized to 1 bit which is available here: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF.
Now, depending on how much RAM you have, you can offload some of the inference to the RAM, but this is a performance bottle neck (less tokens per second)
I haven't tried the 1 bit quantization yet, but I often run the Q4 quantization on a 3090 machine I have Qwen3.8-27B-UD-Q4_K_XL.gguf which is 17.6GB and fits nicely on a 3090 with 24GB of VRAM.
To run these models you need to download them (usually from hugging face) and run them with something like llama.cpp (I personally recommend this over ollama). If you are doing that you need a GGUF file which is available from many people on hugging face, most famously the account named "unsloth" which takes the Safetensors weights that a company publishes and turns it into these GGUF files which can be run on a desktop with llama.cpp.
With the 5060 you are limited. If you can get a 3090 you can do really great local work with Q4 Qwen3.8-27B
In my personal set up (which I will release fully open source soon), my qwen model can actually browse the internet as if it was me (you can actually watch it view pages, its pretty cool), which means your little local model can scrape up to date info off the internet.
Hopefully that was helpful. Open source is the future!