RUN THIS WORKFLOW NOW ON FLOYO!
ABOUT THE WORKFLOW
Read an Image and Answer
Upload an image and type a question or instruction. The model reads what is in the image and answers in plain language. You can also add an audio clip to transcribe or describe alongside the image.
Model
Gemma 4 E4B by Google DeepMind. An open-weights multimodal model built from Gemini 3 research that takes text, image, and audio and writes a text response. Strong at description, analysis, transcription, and question answering, with a configurable thinking mode.