Caption Generator
Generates textual descriptions of objects in a ROI, or of a ROI in the frame.
Overview
The Caption Generator node generates textual descriptions of objects within a Region of Interest (ROI), or of a ROI within the frame. This functionality is useful for applications requiring detailed descriptions of visual elements, enhancing accessibility, and providing contextual understanding.
Inputs & Outputs
- Inputs: 1, Media Format: Raw Video
- Outputs: 1, Media Format: Raw Video
- Output Metadata: Captions
Properties
| Property | Description | Type | Default | Required |
|---|---|---|---|---|
roi_labels | Regions of interest labels | string | — | No |
rois | Regions of interest. Conditional on roi_labels. Format: comma-separated normalized x,y coordinate pairs; separate multiple polygons with semicolons (for example, 0.1,0.1,0.9,0.1,0.9,0.9). | string | null | No |
processing_mode | Processing mode. Options: ROIs, at an Interval (rois_interval); ROIs, upon a Trigger (rois_trigger); Objects in an ROI (objects). | enum | rois_interval | Yes |
trigger | Queue ROI for lookup when this condition evaluates to true. Conditional on processing_mode being rois_trigger. | trigger-condition | null | No |
objects_to_process | ex. car,person,car.red. Conditional on processing_mode being objects. | model-labels | null | No |
tracking_mode | Tracking mode. Options: Centroid (centroid); Top center (top-center); Bottom center (bottom-center); Left center (left-center); Right center (right-center). Conditional on processing_mode being objects. | enum | centroid | No |
min_obj_size_pixels | Min. width and height of an object. Conditional on processing_mode being objects. | number | 64 | No |
obj_lookup_size_change_threshold | Object size ratio change threshold. Conditional on processing_mode being objects. Range: minimum 0.1, maximum 2.0. Step: 0.2. | float | 0.1 | No |
max_lookups_per_obj | Max. attempts per object. Conditional on processing_mode being objects. | number | 5 | No |
interval | Queue objects or ROIs for lookup at least this many seconds apart. Unit: seconds. | number | 1 | No |
display_roi | Display ROI on video? | bool | true | No |
display_objinfo | Display results on video? Options: Disabled (disabled); Bottom left (bottom_left); Bottom right (bottom_right); Top left (top_left); Top right (top_right). | enum | bottom_left | No |
debug | Log debugging information? | bool | false | No |
Metadata
Output Metadata
The fields below are declared by this node's metadata schema; the JSON values are representative examples.
| Path | Type | Description |
|---|---|---|
nodes.<node_id>.rois.<roi_label>.label_changed_delta | boolean | When the caption of an ROI changes |
nodes.<node_id>.rois.<roi_label>.label_available | boolean | Boolean indicating whether the node has a current result for this ROI. |
nodes.<node_id>.rois.<roi_label>.label | string | Current model-generated result for this ROI. |
nodes.<node_id>.recognized_obj_count | integer | Number of objects successfully processed in the current frame. |
nodes.<node_id>.recognized_obj_delta | integer | When one or more new objects are captioned |
nodes.<node_id>.label_changed_obj_delta | integer | When the caption of one or more objects changes |
nodes.<node_id>.unrecognized_obj_count | integer | Number of objects that could not be processed in the current frame. |
nodes.<node_id>.unrecognized_obj_delta | integer | When one or more objects fail to be captioned |
nodes.<node_id>.recognized_obj_ids | array | Array of tracking IDs of objects successfully processed by the node. |
nodes.<node_id>.objects_of_interest_keys | array | Array of metadata keys that contain object IDs relevant to downstream integrations. |
nodes.<node_id>.type | string | Identifies the node type that produced this metadata. |
JSON example
{
"nodes": {
"caption_generator1": {
"label_changed_obj_delta": 0,
"objects_of_interest_keys": [],
"recognized_obj_count": 0,
"recognized_obj_delta": 0,
"recognized_obj_ids": [],
"rois": {
"roi1": {
"label": "example",
"label_available": false,
"label_changed_delta": false
}
},
"type": "caption_generator",
"unrecognized_obj_count": 0,
"unrecognized_obj_delta": 0
}
}
}Object labels and attributes
- Object labels/classes added: In ROI mode, the configured ROI label is added as an object with class
10700. - Object attribute labels/classes added: ROI objects receive
caption_generator_roiwith class10700; generated caption text uses class10701on ROI or upstream objects.
Updated 14 days ago
Did this page help you?
