Tasks
Image generation
Generate images from a description, edit them in plain language, cut out backgrounds, and mark up exactly the part you want changed.
Ask for a picture, an illustration, a mockup, a diagram-like graphic, or a photo and the assistant generates it into the task. Images land on the task’s media board, where you can arrange them, compare versions, and mark them up.
Where it runs: in a task. AI staff can also generate an image during an unattended run, saving it to your Drive.
What the assistant can do
| It can | What that looks like |
|---|---|
| Generate an image | A new image from a text prompt, at the aspect ratio you chose |
| Generate options | Up to four variations of the same prompt, when you ask for choices |
| Edit an image | A new version from an instruction — “warmer light”, “remove the car” |
| Remove the background | The subject isolated on transparency, as a new version |
You get one image unless you ask for options. That is deliberate: a grid of near-identical pictures is rarely what someone wanted.
Aspect ratio is yours, not the model’s
Pick the ratio in the composer — square, portrait, story, landscape, or widescreen — and it is applied. Your choice overrides whatever shape the model would have picked on its own, so a story-format request does not quietly come back landscape.
Editing by pointing
Editing is conversational: describe the change and you get a new version, with the old one kept. When words are clumsy — “this bit here” — drop a comment pin on the image and write the instruction there. The pin travels with the request, so the edit is anchored to the exact spot rather than to your best description of it.
Every edit is a version. Flip back through them, and drag an older version out to restore it.

The original, the edited version with a comment pin on it, and the prompt bar that drives the next edit.
The media board
Images and videos share one board. That is because they are the same material at different stages: any image on the board can be animated into a video, and the result lands beside its source rather than in a different tab. Comment pins are image-only — a video has no fixed frame to pin to.
Getting good images
Describe the picture, not the request. “A ceramic mug on a linen cloth, morning window light, shallow depth of field” beats “a nice product photo”. The assistant expands terse prompts for you, but it cannot guess a specific look you had in mind.
Edit rather than regenerate. An edit keeps everything you liked. A fresh generation rolls the dice on all of it.
Say what to keep. “Same composition, change the background to a studio grey” tells the model what is not up for negotiation.