Deni.AI Labs100% in-browser · no uploads
LIVEBackground Remover · ~26 MB Image to Text (OCR) · ~10 MB Speech to Text · ~40 MB Object Detector · ~170 MB AI Image Describer · ~250 MB Text Summarizer · ~300 MB Sentiment Analyzer · ~70 MB Text to Speech · 0 MB
AI Labs › AI Tools › AI Image Describer

AI Image Describer

Get a short, plain-English sentence describing what appears in a photo, generated entirely on your own device.

VisionModel ViT-GPT2 captioningDownload ~250 MBSpeed 2–6 sPrivacy stays on your device
Drop, paste or click to choose an image
Ready. The model downloads the first time you run it (~250 MB).

Give this tool a picture and it writes a one-sentence description of it, for example "a dog sitting on a couch next to a window". It is a quick way to draft alternative text for images on a website, label a folder of photos, or simply see how a computer vision model interprets a scene.

The image is analysed by a model that runs inside your browser. Nothing is uploaded, so personal photos stay on your machine. The trade-off is that the first visit requires a fairly large download before the tool can start.

How to use it

  1. Select or drag in a JPG, PNG or WebP image.
  2. On first use, allow the model to download; it is roughly 250 MB and is cached by the browser for later visits.
  3. Press the describe button and wait a few seconds while the caption is generated.
  4. Copy the caption and edit it to add details only you know, such as names or places.

How it works

Captions come from ViT-GPT2 image captioning (Xenova/vit-gpt2-image-captioning), loaded with Transformers.js. The model has two halves. A Vision Transformer splits your picture into small square patches and turns them into a numerical summary of what is visible. A GPT-2 language model then reads that summary and writes a sentence word by word, choosing at each step the word most likely to follow. It learned this pairing from a large collection of photos that had human-written captions.

Good for

Limitations

FAQ

Can I use the captions as alt text directly?

They are a good starting point, but good alt text reflects why the image is on the page. Review each caption and add context the model cannot know.

Why is the download so large?

The model combines a vision network and a language network. Once downloaded, it is kept in your browser cache, so you will not need to fetch it again on the same device.

Are my photos sent to a server?

No. The picture is decoded and analysed locally in this browser tab and is never transmitted.

More tools

1

Background Remover

Cut the subject out of a photo and save it as a transparent PNG, processed entirely on your own device.

VisionModel MODNetDownload ~26 MBIn JPG / PNG / WebPOut Transparent PNGSpeed 1–5 s
2

Image to Text (OCR)

Turn screenshots, scanned pages and photos of printed text into editable text without sending the image anywhere.

VisionModel Tesseract (English and more)Download ~10 MBIn Screenshot, scan, photoOut Plain textSpeed 2–10 s
3

Object Detector

Find common objects in a photo and draw labeled boxes around them, using a detection model that runs locally.

VisionModel DETR ResNet-50Download ~170 MBIn JPG / PNG / WebPOut Labeled boxes, countsSpeed 3–10 s
4

Speech to Text

Transcribe English speech from an audio file or your microphone using a Whisper model that runs on your own device.

AudioModel Whisper tiny.enDownload ~40 MBIn MP3 / WAV / M4A / micOut Transcript, timestampsSpeed ~1× realtime
5

Text Summarizer

Condense a long English article or report into a few sentences, generated by a model that runs on your device.

TextModel DistilBART CNN 6-6Download ~300 MBIn English text, 60+ wordsOut Short summarySpeed 10–60 s
6

Sentiment Analyzer

Paste English text and see whether each line reads as positive or negative, with a confidence score for each.

TextModel DistilBERT SST-2Download ~70 MBIn Text, one item per lineOut Positive / negative + scoreSpeed <1 s per line