Nano Banana Playground vs Speech To Markdown: Features, Pricing & Which Is Better (2026)
A side-by-side comparison of Nano Banana Playground and Speech To Markdown — features, pricing, and ideal use cases — to help you decide which AI tool fits your workflow.
Nano Banana Playground
Img Gen Playground (powered by Vercel AI Gateway)
Web-based multi-model image playground for text-to-image generation and image editing with 30+ models via Vercel AI Gateway.
Key features
- Multi-Model Access: Provides unified access to 30+ image generation and editing models from providers such as Google Gemini, Imagen, OpenAI GPT Image, FLUX, Recraft, Seedream, xAI, and ByteDance via the Vercel AI Gateway.
- Text-to-Image Generation: Create images from textual prompts across multiple backends, enabling users to compare stylistic and fidelity differences between models.
- Image Editing: Supports image editing workflows (inpainting/edits) alongside text-to-image generation, allowing iterative refinement of visuals within the same interface.
- Provider Agnostic Interface: Abstracts provider-specific APIs into a single playground so users can switch models and providers without separate integrations or accounts.
- Rapid Model Comparison: Streamlines side-by-side experimentation to evaluate output quality, style, and prompt sensitivity across different model families.
- Built on Vercel AI Gateway: Uses Vercel's AI Gateway and AI SDK for backend connectivity and hosting, simplifying deployment and access to provider endpoints.
- Unified web interface for generating images from text prompts
- Image editing capabilities (in-browser edit flows)
- Support for 30+ models and providers (e.g., Google Gemini, Imagen, OpenAI GPT Image, FLUX, Recraft, Seedream, xAI, ByteDance)
- Powered by Vercel AI Gateway to route requests to multiple model backends
- Built using the AI SDK for standardized integration with model providers
- Model selection and switching within the same playground for comparison and testing
- Export/download of generated assets via the web UI (inferred typical capability)
Best for
- Creative Concepting: Rapidly generate multiple visual concepts for characters, scenes, or product mockups by switching between model backends to explore diverse styles.
- Model Evaluation: Compare output quality and behavior of competing image models (e.g., Google Gemini vs. OpenAI GPT Image) for research or procurement decisions.
- Iterative Image Editing: Upload an image and apply edits or inpainting across different models to refine a visual asset without leaving the playground.
- Marketing Asset Prototyping: Quickly produce and iterate on marketing visuals, banners, or social assets using different model styles to find the best fit.
- Developer Prototyping: Prototype integration flows and prompts for image-generation features before committing to a specific provider or API.
- Educational Demos: Demonstrate differences in generative model capabilities and prompt engineering techniques in workshops or classroom settings.
- Prompt engineering and rapid experimentation across multiple image models
- Comparing output quality and style between different generative providers
- Creating concept art and visual assets via text-to-image generation
- Performing in-browser image edits using different model editing pipelines
- Prototyping integrations that require multi-model image generation access
Speech To Markdown
xajik
Free, 100% local macOS menu-bar app that turns speech into structured markdown using whisper.cpp and any local LLM.
Key features
- 100% Local Pipeline: Runs whisper.cpp for speech-to-text and any local LLM server for structuring — no cloud calls and no API keys required.
- Global Dictation Hotkey: Press ⌘⌥] in any app to have the transcript typed straight at your cursor, works in Terminal, browser, Slack, and more.
- Agent Mode Live Structuring: A floating capsule streams your voice through the LLM into a real-time Markdown, plain text, or HTML document.
- One-Line Install: A single curl-piped script installs xcodegen, whisper-cpp, and ffmpeg via Homebrew, then builds the app from source into /Applications.
- iOS Companion: A fully offline iPhone/iPad app that uses Apple Intelligence on iOS 26+ (iPhone 15 Pro and up).
- Multiple Output Formats: Format, edit, or append the LLM output as Markdown, plain text, or HTML from a single control panel.
- Send-Now Flush: The Send (⏎) control flushes the current buffer to the LLM immediately instead of waiting for the pause/word-count threshold.
- Model Picker: Download and swap Whisper models from Settings — Base (~150 MB) is a good starting point.
Best for
- Private Meeting Notes: Dictate meeting recaps on a Mac with sensitive content that must never leave the device.
- Voice-Driven Coding Comments: Speak function docstrings or PR descriptions into your editor at the cursor via Global Dictation.
- Structured Journaling: Use Agent Mode to ramble freely and get a clean, headed Markdown document out in real time.
- Offline Field Notes on iOS: Capture voice notes on an iPhone with no signal, structured into markdown using on-device Apple Intelligence.
- Slack / Email Long-Form: Dictate long replies straight into Slack or Mail without opening a separate transcription tool.
