Most AI writing tools still treat visuals as someone else’s problem. Google just decided that’s no longer acceptable. As reported by Chrome Unboxed, Gemini can now generate, edit, and refine custom images, diagrams, and infographics directly inside Google Docs, without leaving the document or opening a separate tool.
This is context-aware generation, meaning Gemini reads what’s already in your document and produces visuals that are actually relevant to the content. That’s the part worth paying attention to. It’s not a generic image generator bolted onto a sidebar. The model is pulling from the surrounding text to inform what it creates, which is a meaningfully different interaction model than asking an AI to generate something from scratch with a blank prompt.
For developers and founders who live in Google Workspace, this removes a real friction point. The current workflow, writing in Docs then jumping to Canva, Lucidchart, or even ChatGPT to generate a diagram, then importing it back, is clunky. It also breaks focus. Having generation happen in place, with the ability to refine through follow-up prompts, is a practical improvement to how documents actually get made.
The competitive context matters here. Microsoft has been pushing Copilot deep into Word and the broader Office suite, and it also supports image generation inside documents through integration with Designer and DALL-E. Google is clearly trying to close that gap, but the context-awareness angle is where Gemini could have a real edge if it works as described. Copilot’s image features inside Word have been inconsistent in practice.
The features now available include:
- Custom image generation from text prompts within Docs
- Diagram and infographic creation tied to document context
- In-document editing and iterative refinement of visuals
- Direct insertion without needing external design tools
Zoom out and this is part of a broader push by Google to make Workspace the default environment for AI-assisted knowledge work, not just AI-assisted writing. Adding visual output to that loop is a logical step. The question is execution. Diagram quality, prompt responsiveness, and how well the context-awareness actually holds up in complex documents will determine whether people use this regularly or treat it as a novelty.
Still, the direction is clear. Google wants Gemini to handle more of the document creation process end to end. For anyone evaluating Workspace versus Microsoft 365 for their team right now, this update is worth factoring in.




