GPT-5.6 Sol delivered the strongest overall results in a drawing experiment that asked four AI models to create images stroke by stroke on a digital canvas. TryAI reports in its drawing-arena analysis that GPT-5.6 Sol, Claude Fable 5, Grok 4.5 and Gemini 3.6 Flash received the same colored-pencil tools and either copied famous paintings or worked from text prompts.
The test included reproductions of Leonardo da Vinci’s Mona Lisa and Vincent van Gogh’s Starry Night, plus five original prompts such as a fisherman’s portrait, a rose in a vase and a cabin interior. Models could choose colors, pencil width and pressure, then draw, blend, erase and inspect their canvas during the task.
Gemini 3.6 Flash achieved the highest structural-similarity scores on both reference images. Its final Mona Lisa score was 0.337 on the SSIM scale, ahead of GPT-5.6 Sol at 0.325. Gemini also reached the highest single score in the test, 0.449, during its Mona Lisa run. However, the source notes that this metric measures pixel and structural resemblance rather than perceived artistic quality.
TryAI’s editors rated GPT-5.6 Sol as the strongest visual performer overall. They highlighted its Starry Night and rose drawings, and found it produced more detail than its rivals in most tasks. Claude Fable 5 placed second in their subjective assessment, while Grok 4.5 produced the weakest results.
More review did not mean a better image
A notable pattern emerged in all eight reference-image runs: every model ended with a lower score than it had achieved earlier. Gemini reviewed its canvas about 23 times per drawing, yet its Mona Lisa score fell from its peak of 0.449 to 0.337 at the end. GPT-5.6 Sol similarly dropped from 0.352 to 0.325.
Cost also varied sharply. TryAI estimates Claude Fable 5 cost $160.58 for seven drawings, compared with $7.74 for GPT-5.6 Sol, $9.21 for Grok 4.5 and $12.87 for Gemini 3.6 Flash.
Stay up to date
AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox: