闲社服务运行正常AI智能体自动化平台
g

图像定位

技能标识:grounding

Use GLM-4.7V's multimodal grounding capability to detect and locate objects/text in images. Activate when user asks to find, locate, detect, or ground specific objects, text, UI elements, or regions in an image. Also triggers on phrases like "找到xxx的位置", "框出xxx", "定位xxx", "grounding", "bounding box", "坐标框".

作者:admin | 来源记录:ClawHub
登记来源
ClawHub
版本
V 1.0.0
检测标记
后台标记通过
免费
免费获取
0
收藏
来源与检测标记为平台登记信息,并不代表已展示可复核的检测报告。使用前请核对版本、依赖和所需权限,建议先在隔离环境中运行。
概述
安装方式
版本历史

图像定位

Grounding - 多模态目标定位

利用 GLM-4.7V 的 grounding 能力,在图片中定位目标对象或文字,输出带标注框的结果图。

工作流程

用户输入(图片 + prompt)
│
▼
HttpInterface() → 调用模型 API → 得到 response 文本
│
▼
parsebboxesfrom_response() → 从回复中解析出坐标框列表
│
▼
visualize_boxes(renormalize=True) → 反归一化 + 画框 → 保存结果图

Step 1: 调用模型获取坐标

使用 HttpInterface 调用模型 API:

python
import os
os.environ[NO_PROXY] = # 跳过代理
os.environ[no_proxy] =

from interface_http import HttpInterface

url = http://:/v1/chat/completions
prompt = 请在这张图中找到所有{target},并以 [xmin, ymin, xmax, ymax] 格式输出每个目标的边界框坐标,坐标值为 0-1000 的归一化整数。每个目标一行,格式如下:
目标名称: [xmin, ymin, xmax, ymax]

response = HttpInterface(url, prompt, images=[imagepath], nothink=True)

返回: 目标名称: [xmin, ymin, xmax, ymax]

注意: 调用前需设置 NO_PROXY 环境变量跳过代理,否则内网请求会被代理拦截。

Step 2: 解析坐标框

python
from utilsboxes import parsebboxesfromresponse

boxes = parsebboxesfrom_response(response)

返回: [[x1, y1, x2, y2], ...] (0-1000 归一化)

parsebboxesfrom_response 会自动:

  • - 从回复尾部向前检查截断,拓展 context window
  • 遍历所有括号风格([], {}, (), <>, )提取坐标
  • 扁平化嵌套列表,返回一维 box 列表

Step 3: 画框可视化

python
from utilsboxes import visualizeboxes

visualize_boxes(
imgpath=imagepath,
boxes=boxes, # parsebboxesfrom_response 的输出
labels=[label1, label2], # 每个框的标签
renormalize=True, # 自动将 0-1000 归一化转为像素坐标
save_path=output.jpg,
colors=[red, blue], # 可选
thickness=[2, 3], # 可选
)

renormalize=True 时,内部自动调用 reversenormalizebox:pixel = coord * img_dimension / 1000

完整示例

python
import os
os.environ[NO_PROXY] = 172.20.112.202
os.environ[no_proxy] = 172.20.112.202

from interface_http import HttpInterface
from utilsboxes import parsebboxesfromresponse, visualize_boxes

url = http://172.20.112.202:5002/v1/chat/completions
img = /path/to/image.jpg

1. 调用模型

response = HttpInterface( url, 请在这张图中找到红色圣诞帽,以 [xmin, ymin, xmax, ymax] 格式输出坐标(0-1000归一化), images=[img], no_think=True, )

2. 解析坐标

boxes = parsebboxesfrom_response(response)

3. 画框

visualizeboxes(imgpath=img, boxes=boxes, labels=[圣诞帽], renormalize=True, save_path=out.jpg)

工具函数速查

函数作用
HttpInterface(url, prompt, images, nothink)调用模型 API,返回文本回复
parsebboxesfromresponse(text)
从模型回复中提取所有坐标框列表 | | findboxesall(text, flat=True) | 提取文本中所有括号风格的坐标框 | | reversenormalizebox(box, w, h) | 0-1000 归一化 → 像素坐标 | | visualize_boxes(..., renormalize=True) | 画框 + 自动反归一化 |

注意事项

  • - 模型 API 地址配置在 /root/.openclaw/agents/main/agent/models.json
  • 调用内网模型时必须设置 NOPROXY 环境变量
  • nothink=True 可关闭模型思考模式,加快响应

标签

skill ai

通过对话安装

以下为平台配置的接入选项,并非逐项实测通过。能否安装取决于客户端支持、技能来源和运行环境:

OpenClaw WorkBuddy QClaw Kimi Claude

方式一:安装 SkillHub 和技能

帮我安装 SkillHub 和 visual-grounding-1776115683 技能

方式二:设置 SkillHub 为优先技能安装源

设置 SkillHub 为我的优先技能安装源,然后帮我安装 visual-grounding-1776115683 技能

通过命令行安装

skillhub install visual-grounding-1776115683

下载

⬇ 下载 grounding v1.0.0(免费)

文件大小: 648.48 KB | 发布时间: 2026-4-17 16:29

v1.0.0 最新 2026-4-17 16:29
Initial release of visual-grounding skill using GLM-4.7V's multimodal capability:

- Supports detection and localization of objects, text, and regions in images, with bounding box output.
- Activates automatically on user prompts related to finding, locating, or grounding visual elements.
- Provides step-by-step workflow: model API call, response parsing for bounding boxes, and result visualization with labeled boxes.
- Includes utility functions for API interaction, coordinate parsing, normalization, and image annotation.
- Offers quick reference and usage examples for easy integration.
返回顶部