Ask AI about a screenshot
Describing a visual problem in words is slow and lossy. Capturing it takes two seconds. SideNote Pro has capture built in so the screenshot goes straight into the conversation.
Choosing a capture mode
| Mode | Use it for |
|---|---|
| Region | One dialog, one chart, one paragraph. The mode you will use most. |
| Active window | A whole application window, without cropping by hand |
| Current screen | Everything on one monitor, including how things sit together |
| Entire desktop | The full virtual desktop across every display |
Region capture hides SideNote Pro, shows the desktop as a backdrop with a translucent overlay, and lets you drag out the area you want. Esc or a right-click cancels. Active-window capture uses the window that was in front before you summoned the assistant, so it captures what you were looking at rather than the assistant itself.
Prefer region capture where you can. A tighter crop means the model spends its attention on the thing you are asking about instead of on your taskbar.
Other ways to get an image in
- `Ctrl+V` pastes a bitmap from the clipboard. This is the bridge to whatever snipping tool you already use.
- Drag and drop an image file onto the window.
- The image picker for PNG, JPEG or JPG, and WebP.
Up to five images per request, previewed before sending, removable individually, and kept in a deterministic order so *the second screenshot* means what you expect.
Asking a good question about an image
An attached image with no question attached tends to produce a description of the image, which you did not need since you were looking at it. Say what you want.
An application error
*This dialog appears when I save. What is it actually telling me, and what are the three most likely causes?* Add the context the screenshot cannot contain: what you did just before, and what you expected.
A chart
*What does this chart show, and what would you check before trusting the trend in the last quarter?* The second half is the useful half, because a description of a chart is rarely what you wanted.
A UI problem
*This layout breaks at this width. Which element is most likely causing the overflow?* Pair it with local folder context over your source and you get a specific file rather than general advice.
Part of a web page
*Summarize the argument on this page and list the claims made without evidence.* Region capture is better here than pasting text, because it keeps the layout that tells you what is emphasis and what is a footnote.
A design review
*Review this screen for usability problems. Focus on whether a first-time user could work out what to do next.* Naming the lens gets you a review rather than an inventory of elements.
Your model has to support images
Image understanding is a model capability. SideNote Pro sends images in the OpenAI-compatible multimodal shape to any compatible endpoint, but a text-only model will ignore or reject them. If a model seems to be answering without having looked, check that it is a vision-capable model.
This is worth checking twice on local providers: many small local models are text-only. And SideNote Pro performs no OCR itself, so whatever text is read out of your screenshot is read by the model's own vision capability.
Capture is always something you started
There is no background capture, no periodic screenshotting and no ambient screen awareness. A screenshot exists because you started one, and it is transmitted only when you send the message it is attached to.
One practical note: protected or DRM-restricted content may capture as blank. Windows blocks it at the system level, which is not something an application can work around.