I sometimes have things to discuss with LLMs, which are more private than usual. I was thinking how to do it in a way that doesn't reveal my identity.
1. Use Tor to access the provider?
2. Create a random account?
3. Use some form of untraceable payment (which one?)
4. Scrub all information provided to the LLM from personally identifiable information?
It seems like a lot of effort. So is running a local LLM, for which I don't even have the hardware. How do you do it?
- llama-server -hf ggml-org/gemma-4-26b-a4b-it-GGUF:Q4_K_M
After that simply open browser and enter: http://localhost:8080
What this does: This will download Gemma4 AI with 26B param & start a http server for chat
Its shockingly capable for its size. Does it beat the top end models? No, but as long your don't fall into the hallucations. Its just fine.
Edit: the software is llama.cpp you can download it from "releases" which u can find at github right side. No need to know how to build it
Edit2: Pro tip is, only use chat per context you want to use. Lots users want "dynamically" change the context, but that doesnt really work from my experience.