Frequently asked questions
Speakybara is a free Windows app that turns any part of your screen into text you can copy or listen to: press a hotkey, drag a frame, done.
Everything people ask before and after installing, grouped by subject. If your question is not here, write to us and it probably ends up on this page.
Ask something elseGetting started
Speakybara is a free Windows app that turns any part of your screen into text you can copy or listen to: press a hotkey, drag a frame, done. It works on anything that is visible, including text you cannot select, such as a scanned PDF, an error dialog or a screenshot a colleague pasted in Teams. The recognition runs on your own PC with the OCR engine that is built into Windows.
Yes, and it stays free, because the expensive parts are optional. Recognition happens on your own PC and the local Piper voices run on your own processor, so using Speakybara costs us nothing per user. The paid plan, Speakybara Managed at €59 a year, unlocks no features at all: it only removes the Google Cloud setup.
No. There is no sign-up, no email address, no trial period and no card. Download the zip, unzip it, start Speakybara.exe and the app works. An account only exists if you buy Speakybara Managed, because a licence key has to belong to someone.
Windows 10 build 19041 and newer, and Windows 11, both 64-bit. It installs for your own user account, so Windows does not ask for an administrator password. The full list of requirements is on the download page.
Download the zip, right-click it and choose Extract All, then double-click Speakybara.exe in the folder; it takes about two minutes and nothing gets installed. Windows SmartScreen will warn about an unknown publisher because the app is not signed yet, so choose More info and then Run anyway. Speakybara then opens its window and puts an icon next to the clock.
Yes. About 55 MB of memory while it is reading, and 0 MB of video memory, because the graphics card is never used. It sits in the tray doing nothing until you press a hotkey, and recognising a normal selection takes 50-150 ms, so an ordinary or older laptop is enough. The exception is a local AI model you choose to run yourself, such as Ollama or LM Studio: that is separate software with its own memory and often graphics-card needs, outside what Speakybara uses.
A macOS version is planned, but there is no date and no waiting list to sign up for. Speakybara is a Windows app today because it leans on the OCR engine and the global hotkeys that Windows provides. We would rather say that plainly than collect email addresses for a date we cannot promise.
Shooting, copying and listening
Press Ctrl+Alt+E to copy or Ctrl+Alt+S to listen, drag a frame around the part of the screen you mean, and let go. Speakybara recognises the text inside that frame on your PC in 50-150 ms and either puts it on your clipboard or starts reading it aloud in the read-along window. There is nothing else to click.
Yes. Ctrl+Alt+E puts the text inside the frame straight on your clipboard and speaks nothing. It uses the same recognition step, so it works on scanned pages, screenshots and dialogs where selecting text is impossible. Paste it wherever you would otherwise have typed it over by hand.
Never guaranteed to be perfect: a blurry screenshot, a very thin font or a table can come out wrong. For anything that has to be exact — amounts, codes, IBANs, addresses — select the text and press Ctrl+C wherever the text is selectable, instead of trusting a capture. A small clean-up always runs on the result: broken lines are rejoined, and stray page numbers and bullet marks are dropped. The extra AI clean-up under Settings › Text recognition is newer and still being improved, which is why it sits on the roadmap rather than in the finished list.
Yes, the frame you drag is an ordinary screenshot and you can keep it as an image file. It is a convenience rather than the point: Speakybara exists for the text inside the picture. There is no library, no annotation and no upload.
No. Recognised text is kept in a local history you can switch off, but there is no browsable library of full screenshots — only a small preview next to each history row, which you can switch off — no annotation tools, no uploads and no share links. Speakybara does the step those tools skip: turning what is in the picture into text you can copy or hear.
No. A screen reader such as NVDA or Windows Narrator describes a whole interface for people who cannot see the screen. Speakybara is a read-aloud tool: you point at one piece of screen and it reads that piece. It can sit next to a screen reader usefully, because it reaches text inside images that a screen reader cannot.
Yes, because it reads pixels rather than the structure of a document. Browsers, Outlook, Word, PDF viewers, Teams, Citrix and remote desktop sessions, video players, installers, error dialogs: if it is on your screen you can put a frame around it. The only thing it cannot reach is text that is not visible, such as the part of a page you have not scrolled to.
That is the main reason people install it. A scanned PDF, an image in an email, a screenshot pasted in Teams, an error message with no copy button, a dashboard label, a paused video frame: to Speakybara they are all just pixels. Recognition takes 50-150 ms and happens on your own PC.
Yes, press Ctrl+Alt+C and Speakybara reads the current clipboard contents aloud. No frame, no selection. It is useful when you have already copied something and would rather hear it than read it.
Yes, every one of them. Right-click the tray icon, open settings, go to shortcuts, click a field and press the combination you want. If another program already owns that combination, Speakybara says so instead of quietly failing.
Yes. Speed runs from 0.8x to 1.75x, and pause, stop and repeat each have a hotkey. In the read-along window the sentence being spoken is highlighted and its words turn white as they are said, and you can click any sentence to continue from there.
Recognition takes 50-150 ms for a normal selection, so speaking starts almost as soon as you release the mouse. Repeated sentences come from a cache on your own PC, which makes re-reading instant and free. We make no claims about how much faster listening makes you; that depends entirely on you.
Accessibility
It is one of the reasons people install it. Listening moves the effort from decoding the letters to understanding the sentence, and hearing and seeing the same sentence at once gives your attention two things to hold on to. It is not a treatment, and it does not replace NVDA or Narrator — it works alongside them, on the text they cannot reach. There is more on the ADHD page and in the reading section of Features.
For a lot of people it does: hearing and seeing the same sentence gives your attention two things to hold on to, and part of the effort shifts from decoding the words to understanding them. It is not the same for everyone. On easy text you already read quickly, the audio can be more distraction than help, so it tends to matter most on the pages that are slow or dense for you. Try it on one hard page before you decide.
Voices
Three services, switchable in settings at any time. Piper runs locally on your processor, free and offline. Microsoft Edge neural voices are free and need internet. Google Cloud voices sound best and need either your own API key or Speakybara Managed.
Yes, with the Piper voices, which run on your own processor. Text recognition is offline in every case, so copying text always works without a connection. The Edge and Google voices need internet because the sentence text is sent to Microsoft or Google to be spoken.
Piper is an open-source speech engine that runs locally, and Speakybara ships Dutch and English Piper voices. They sound noticeably more synthetic than the cloud voices, and they never send anything anywhere. Each extra voice is about 60 MB on disk and uses 0 MB of video memory.
Yes, the same neural voices Edge uses to read pages aloud, and they cost nothing. You need no account and no key. The sentence text is sent to Microsoft to be spoken, so they only work with an internet connection.
Yes, and for most personal use it is free: Google's own free tier is 4 million characters a month for Standard and WaveNet voices and 1 million for Neural2 and Chirp 3 HD. You create a Google Cloud project, enable Text-to-Speech, make a key and paste it into settings — about 15 minutes of work. Speakybara Managed exists purely for people who would rather not do that.
Yes, both, and it works out which is which by itself — per selection, per sentence and per word. That way English terms in a Dutch email such as 'pull request' or 'deadline' sound English instead of being mangled, while Dutch look-alikes like brand, post and model stay Dutch. You can add your own terms to a word list.
Privacy
No. The frame you drag is recognised on your own PC by the OCR engine built into Windows and the image is then discarded; the recognised text stays behind in a local history you can switch off, and that file never leaves your machine. Only the sentence being read goes out — to Microsoft or Google — if you pick one of their voices.
It depends on the voice you choose. With Piper nothing leaves your PC at all. With Microsoft Edge voices the sentence text is sent to Microsoft to be spoken, and with Google voices — through your own key or through Speakybara Managed — the sentence text is sent to Google. The image is never sent in any of these cases.
The image is discarded straight after recognition, but the recognised text is kept in a local history you can switch off, and that file never leaves your PC. Sentences spoken by the online voices are also cached as audio on your own PC so repeating them is instant and costs no characters; that cache holds at most 600 files or 250 MB and you can clear it in settings. With Managed the sentence text passes through our service on its way to Google, and the privacy page states exactly what is kept.
The image never leaves your PC, so the screenshot itself stays local whatever you choose. If the text must not leave the building either, use the copy hotkey, which involves no voice service at all, or pick the Piper voices, which run offline. Check with whoever owns your IT policy before installing anything on a work machine.
Free vs Managed
Nothing to the app, and that is the point. Managed is the same Speakybara: no extra features, no locked buttons, no watermark. What you buy is not having to create a Google Cloud project, an API key and a billing account — roughly 15 minutes of setup you never do. It costs €59 a year or €6.99 a month, VAT included.
No, and we would rather say so here than let you find out later. The voices are the same Google voices; with your own key you pay Google directly and most personal use stays inside their free tier. Managed sells the setup you skip, not better sound, so if you enjoy doing it yourself, do it yourself — that is the normal option, not the second-class one.
400,000 characters a month, roughly 133 A4 pages. The first 60,000 characters, about 20 pages, use Google Chirp 3 HD, the best voice available; after that Speakybara continues automatically on WaveNet for the rest of the calendar month, switching at the next sentence boundary. Repeated sentences come from the cache and cost nothing.
The app keeps working. Nothing is locked, nothing is deleted, no watermark appears: you fall back to the free Piper and Microsoft Edge voices, exactly as before you bought anything. The same happens if you use up the 400,000 characters within a month — Managed pauses until the 1st and the free voices take over.
Yes. Enter your company name and VAT number at checkout and the invoice states the VAT, which is what most expense systems need. In the Netherlands a read-aloud tool is often treated as a compensating aid at work, but whether that applies to you is between you, your employer and UWV, and we cannot promise anything about it.
Licence and devices
Install Speakybara, right-click the tray icon, open settings, go to voice, choose Managed and paste the key. The app checks it once and then about once an hour, and keeps working offline for up to seven days if the check cannot reach us. It takes about a minute in total.
3 at the same time, so a desktop and a laptop both fit. You can free a slot from the app or from your account page whenever a machine goes away. The licence belongs to you, not to a particular PC.
Sign in with the email address you bought it with; the key is on your licence page, where you can copy it or have it emailed again. If you no longer have access to that address, write to support@speakybara.app from the closest one you still have and we will sort it out.
Billing
€59 a year or €6.99 a month, both including VAT. Paid yearly that works out at about €4.92 a month, roughly 30% less than paying monthly. Everything else in Speakybara stays free.
Yes. Every payment produces an invoice stating the VAT rate and amount, downloadable from your account at any time. Add your company name and VAT number at checkout if they need to appear on the document.
Yes, within 14 days, for any reason and with no deductions, even if you already used the Managed voices. That is the EU right of withdrawal and we do not ask you to sign it away at checkout. After that the subscription simply runs to the end of the period you paid for.
Turn renewal off on your account page; it is one click and no email exchange. Managed then runs to the end of the period you paid for and stops, and the app falls back to the free voices. A reminder goes out 30 days before every renewal, so it never arrives as a surprise.
Troubleshooting
That is expected. The app does not carry a code-signing certificate yet, and SmartScreen warns about anything it has not seen often enough. Choose more info and then run anyway; if you want certainty first, compare the SHA-256 on the download page with the file you have.
Usually another program has claimed the same combination: Snipping Tool, a screenshot tool, Teams and game overlays all like these keys. Open settings, go to shortcuts and pick a different combination; Speakybara reports a conflict when it can detect one. Also check that the tray icon is there at all — if it is not, the app is not running.
Windows OCR reads what it can see, so size and contrast decide the result: a slightly wider frame around crisp text beats a tight frame around small, low-contrast text. Zoom the page in before you drag, or raise the upscale factor in settings. For a language other than Dutch or English, the matching Windows language pack has to be installed.
Check the voice service first: the Edge and Google voices need internet, and without it Speakybara falls back to a local voice or stays quiet. Then check the volume slider in settings and the Windows output device. If you are on Managed, the status page shows whether the Managed voice service is having a bad day.
Next to the app itself, in a file called speakybara-error.log. It records what went wrong, not what you read: no text from your screen ends up in it. Attaching it to a support message usually saves a round trip.
Last updated September 9, 2026
