Skip to content

Repository files navigation

Visual Bridge

把甲方脑子里的画面,变成设计师能直接开工的需求。

「高端大气一点」「再有科技感一些」「感觉不太对」—— 这些话每出现一次,设计就多返工一轮。Visual Bridge 做的是中间那层翻译:把模糊的业务黑话,变成具体的视觉要素和可执行的英文提示词。

CleanShot 2026-08-07 at 18 43 13@2x

信息不齐时,它不出图,先把缺的那几项问回来。


它和「AI 生图工具」不一样在哪

它默认不生成。

大部分工具你说一句就出图,出来不对再重说一遍,来回全靠碰运气。Visual Bridge 反过来 —— 信息不全就不生成,词模糊就先跟你对齐。

具体是三条判定:

情况 它怎么做
要素不齐(只说「我要张海报」) 追问缺的那几项,不生成
要素齐了,或你明确说「就这样,生成」 生成
你提了模糊的修改(「再大气一点」) 先说明它打算怎么翻译这个词,等你确认

第三条是这个工具的核心。「大气」不是一个视觉指令,「增加留白 + 黑金配色」才是。它会先问你「您指的『大气』是这个意思吗」,确认了再动手。

省的不是算力,是你和设计师之间那几轮来回

CleanShot 2026-08-07 at 18 48 57@2x

遇到「再大气一点」这类模糊修改,它先说明打算怎么翻译这个词,等你确认。

它收集什么、产出什么

必须齐的四项要素

应用场景   社交媒体海报 / 视频封面 / 舞台背景 / 广告素材
画面主体   要展示的人、物、文字或核心元素
风格氛围   科技感 / 国潮 / 极简 / 温馨 / 赛博朋克……
色调光影   冷暖、明暗的主要倾向

产出

  • 4 组英文提示词 —— 构图、角度或细节上略有差异,给你挑
  • 推荐画面比例 —— 按场景推:社交媒体 1:1、竖版海报与短视频 9:16、横屏与舞台背景 16:9、电商详情页 3:4、演示文稿 4:3
  • 思考链路 —— 每一步用了什么依据、参考了哪个知识库,都摊开给你看
CleanShot 2026-08-07 at 18 47 32@2x CleanShot 2026-08-07 at 18 47 51@2x

出图不是终点。四组结果可以直接基于提示词继续改,也可以拿图生图,快速逼近最优解。

知识库与记忆库:不会写提示词也能用

这一条是给运营伙伴准备的 —— 不需要先学会写提示词,才能把需求说清楚。

  • 知识库 内置各个生图平台的优秀案例和提示词模板,转化的时候直接调用
  • 记忆库 留住这次对齐过的翻译和偏好,下次不用从头再讲一遍

输出里的思考链路会标出每一步用了哪个知识库(knowledgeUsed),所以它引用了什么,你能看见。

CleanShot 2026-08-07 at 18 46 27@2x

黑话映射表(节选)

这张表是整个工具的底子:

高端大气    → minimalist design, premium materials, soft ambient lighting,
              muted color palette, elegant composition
科技感      → futuristic, holographic elements, neon accents,
              dark background, circuit patterns
温馨        → warm lighting, cozy atmosphere, soft textures,
              earth tones, natural materials
年轻活力    → vibrant colors, dynamic composition, bold graphics,
              energetic mood
更有感觉    → cinematic lighting, depth of field,
              emotional storytelling, dramatic shadows

完整映射在 constants.ts。要加自己行业的黑话,改这一处就行。

支持的场景

Social Media Images              社交媒体图片
Video / Short Drama Resources    视频与短剧素材
Advertising Image Resources      广告图片素材
Stage Background Images          舞台背景图

本地运行

环境要求: Node.js v18 或更高

npm install
npm run dev

默认跑在 http://localhost:3000

需要在根目录建一个 .env,填 GEMINI_API_KEY(该文件已被 git 忽略)。

技术栈: React · TypeScript · Vite · Google Gemini API · Cloudflare Worker

部署

推送到 main 触发 .github/workflows/deploy.yml,自动构建并发布到 gh-pages 分支。

首次需要在仓库 Settings → Pages 里,把 Source 设成 Deploy from a branch,选 gh-pages 分支和 / (root) 目录。(gh-pages 分支在第一次 Action 跑成功之后才会出现。)

vite.config.ts 里配了 base: './',所以部署在任何子目录下都能正常工作。


为什么做这个

我本职工作里,最耗人的从来不是设计本身,是**「我说的」和「他理解的」之间那段距离**。

甲方说「高端大气」,设计师听到的是一个没有边界的形容词;做出来第一版,甲方说「不是这个感觉」,但也说不出到底哪里不对。三轮下来,双方都累,而问题从头到尾都不在审美上 —— 在没人把那个词翻译成具体的东西。

所以这个工具最重要的功能不是生图,是拦住那句模糊的话,先把它拆开

做的过程里我发现,让 AI「不要急着给结果」比让它给结果难得多。默认行为是有求必应,你得写很长的规则去拦它,还得给它一套判断标准 —— 什么叫要素齐了、什么叫模糊、模糊的时候该怎么反问。这部分逻辑在 constants.ts 的系统提示词里,比界面代码长。

这和我在另一个项目 Owli 上的判断是同一件事:AI 的产出只能当方向,人来推动结果。 区别只是 Owli 拦的是调研结论,这里拦的是一句「再大气一点」。


English

Visual Bridge

Turns the picture in a client's head into a brief a designer can actually start from.

"Make it more premium." "Give it more of a tech feel." "It's just not quite right." Every one of these costs another round of revisions. Visual Bridge does the translation layer in between — turning vague industry shorthand into concrete visual elements and executable English prompts.

What makes it different from an AI image tool

It doesn't generate by default.

Most tools generate on your first sentence, and if it's wrong you rephrase and roll the dice again. Visual Bridge inverts that: no generation while information is missing, and no generation on a vague edit until you've agreed what it means.

Situation What it does
Requirements incomplete ("I need a poster") Asks for what's missing. No generation.
Requirements complete, or you say "just generate it" Generates
You give a vague edit ("make it more premium") Explains how it plans to translate that word, then waits for your confirmation

That third row is the point. "Premium" isn't a visual instruction; "more whitespace, black-and-gold palette" is. What it saves isn't compute — it's the rounds between you and your designer.

What it collects and what it returns

Four required elements: scenario (social post / video cover / stage backdrop / ad creative), subject, style and mood, colour and lighting.

Output: four English prompt variations (differing in composition, angle or detail), a recommended aspect ratio derived from the scenario (1:1 social, 9:16 vertical poster and short video, 16:9 landscape and stage, 3:4 e-commerce, 4:3 decks), and the full reasoning chain — every step, and which knowledge base it drew on.

Knowledge base and memory: usable without knowing how to prompt

This part is for the marketing and ops people — you should not have to learn prompt engineering before you can state a requirement. A built-in knowledge base carries strong examples and prompt templates from the major image platforms, and a memory store keeps the translations and preferences you already agreed on, so you do not re-explain them next time. Every reasoning step reports which knowledge base it used (knowledgeUsed), so what it drew on stays visible.

The jargon map

高端大气 (premium)  → minimalist design, premium materials, soft ambient lighting…
科技感 (techy)      → futuristic, holographic elements, neon accents…
温馨 (warm)         → warm lighting, cozy atmosphere, soft textures…

The full map lives in constants.ts — add your own industry's shorthand there.

Running locally

Node.js v18+, then npm install and npm run dev (serves on http://localhost:3000). Put GEMINI_API_KEY in a .env at the root — that file is git-ignored.

Built with React, TypeScript, Vite, the Google Gemini API and a Cloudflare Worker. Pushing to main triggers .github/workflows/deploy.yml, which builds and publishes to the gh-pages branch.

Why I built it

In my day job the expensive part was never the design — it was the gap between what I said and what they heard. A client says "premium," a designer hears an adjective with no edges, and three rounds later everyone is tired over a problem that was never about taste. Nobody had translated the word into something concrete.

So the important feature here isn't generation. It's stopping the vague sentence and taking it apart first.

Building it, I found that getting an AI not to answer is much harder than getting it to answer. The default is eager compliance; holding it back takes a long set of rules and a standard for judging what counts as complete, what counts as vague, and how to ask back. That logic — in the system prompt in constants.ts — is longer than the UI code.

It's the same conclusion I reached on Owli: AI output is a direction, not a result; a person drives the outcome. Owli stops a research conclusion. This one stops "make it more premium."

About

Turns a client's vague idea into a brief a designer can actually start from · 少来回三轮

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages