Whistle: Speech to Text in 16.9 MB (cactuscompute.com)
skolos 2 hours ago
stronglikedan an hour ago
I chuckled at this because my inner voice had an accent as I was reading your comment, due to your writing style.
yuchi 41 minutes ago
Nition 30 minutes ago
Zacharias030 28 minutes ago
ASalazarMX 10 minutes ago
schappim 21 minutes ago
skolos 18 minutes ago
mrguyorama 11 minutes ago
With a restricted grammar, built in Windows voice recognition, all on device, has managed this exact use case quite well for over a decade. I used it to try and build a clone of the various paid apps that allow you to issue orders to Arma soldiers with voice commands
INTPenis 4 hours ago
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
ComputerGuru 4 hours ago
Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...
cgbur 4 hours ago
Imustaskforhelp 3 hours ago
Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)
Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.
The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.
I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.
Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.
asa123 2 hours ago
"what are your observations on feeling as if sharing that output though?"
and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."
could you rephrase the question
for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.
dv35z 3 hours ago
You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.
You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.
flockonus an hour ago
I prompt it to:
"Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings. Keep original in parenthesis when not sure if it's a misunderstanding. Do not summarize or alter the nature of the content, simply tidy the transcript."
lnenad 32 minutes ago
johanvts 11 minutes ago
apitman 30 minutes ago
In case it's helpful to anyone else using it, at first it felt a bit slow to me, because there was a noticeable pause after I finished a message before it would quickly type it all out. I changed the input method from direct to clipboard and it's way faster now, almost instantaneous.
nvtop 4 hours ago
boplicity 4 hours ago
xp84 3 hours ago
hbn 2 hours ago
raddan 2 hours ago
thayne an hour ago
It's also really terrible at recognizing names of my contacts, probably because those names are not represented in the training data.
testycool 4 hours ago
yymir 4 hours ago
yu3zhou4 3 hours ago
I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.
apitman 22 minutes ago
islewis 2 hours ago
The usecase for small models like this is making on-device STT/TTS more accessable. This is important if your usecase is sensitive to either privacy or latency, but this comes at the cost of quality.
My experience has been that these small TTS models are unexpectedly good if your audio is in distribution (western accents, higher quality audio, common vocabulary), but pretty quickly degrade as you move outside of that. They often dont support more complex features such as diarization, multilingual, or realtime streaming either.
Joel_Mckay 2 hours ago
In some cases, this may improve function for a few hours. Best regards =3
albert_e 4 hours ago
iforgotmypasswo 3 hours ago
paynedigital 2 hours ago
solarkraft 3 hours ago
Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!
nicksaroha 2 hours ago
raddan 2 hours ago
zimpenfish 3 hours ago
jwr 3 hours ago
aqfamnzc a minute ago
wkcheng 4 hours ago
This definitely seems lighter and faster. How does accuracy compare?
theturtletalks 3 hours ago
I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.
jwr 3 hours ago
I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).
weitendorf 2 hours ago
If you’re feeding the results into a very smart LLM, it will figure out what you meant (but crucially ONLY if you warn it or tell it to do so, in some cases!). If you’re writing code directly or creating something for public consumption, you can’t tolerate mistakes. If you’re taking notes for yourself you just want it to work cheaply.
If you are ok with the complexity you can run both, and a Meta/Google open model with native audio, and let a smart LLM doctor it up. If performance really matters you can pay for a proprietary model or train one yourself. Until you get to that point, I think it probably doesn't matter much either way. Voice is just too easy to fiddle with
andy_ppp 5 hours ago
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
joewhale 4 hours ago
mejutoco 4 hours ago
jasonwatkinspdx 4 hours ago
stymaar 3 hours ago
[1]: https://fr.wikipedia.org/wiki/Langage_siffl%C3%A9_d%27Aas
charv 4 hours ago
amelius an hour ago
properbrew 2 hours ago
TomGarden 2 hours ago
sfpk 3 hours ago
jakobov 2 hours ago
Lebenita an hour ago
alasano an hour ago
Built my own in a day that blows them out of the water. You can quite literally pick any sufficiently good local model or API provider and combine it with Cerebras for cheap and very fast AI post processing and formatting.
MisterMunchkin 3 hours ago
e12e 4 hours ago
Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).
So, reasonable, but limited?
armcat 4 hours ago
kamranjon 4 hours ago
rafaelm 3 hours ago
rshemet 2 hours ago
opening this thread for questions/feedback if you have any
jayshah5696 4 hours ago
Centigonal 3 hours ago
lab14 an hour ago
mo2art 4 hours ago
jjice 4 hours ago
mrkn1 4 hours ago
mrkn1 4 hours ago
pzo 3 hours ago
saturn8601 4 hours ago
MayeulC 4 hours ago
saturn8601 3 hours ago