HTML Extract Node
Overview
- Type:
transform.extract.html - Display Name: HTML Extract
- Category: transform
- Execution:
supportsExecution: true - I/O Handles:
- Input:
defaultIn(left) - Outputs:
result(Success),error(Error)
- Input:
Purpose: Turn a device's web page into structured JSON. Some scales, printers and legacy controllers expose their state as a small HTML page. The REST connector (or the TCP connector, once shipped) fetches the bytes; this node pulls named fields out with CSS selectors or XPath expressions and emits JSON the rest of the pipeline can consume directly. Framing and encoding remain the caller's job.
The output can be:
- A flat object when every selector returns a single scalar (v1 default).
- An array of scalars when a selector opts into
multiple: true. - An array of objects when a selector opts into
multiple: trueand definesfields:— each match becomes an object, sub-selectors map to per-item keys.
When to use it
- The device exposes an HTML page with the values embedded server-side (visible in the raw response, no JavaScript needed).
- You want to name fields the pipeline reads downstream, so a firmware update that reorders spans does not silently blank the value the way a positional regex would.
- The values you need are text nodes, attribute values, or small HTML fragments.
When not to use it
- The page only becomes populated after JavaScript runs. Use a purpose-built scraper; this node does no JS execution by design.
- The endpoint already returns JSON. Use the REST connector's JSON parsing directly; you do not need this node.
- The device speaks a named protocol (Modbus, OPC UA, EtherNet/IP, MQTT). Use that connector.
Parameters
input (required)
An expression that resolves to an HTML string, for example {{ $input[0].result.body }} or {{ $node["HTTP"].result.body }}. Only expressions are accepted (static text is refused), so the two-mode ambiguity of "path or expression" from other transform nodes does not apply here.
selectors (required, at least one)
Each entry is one output field:
name(required): the key on the output object. Names must be unique — duplicates would silently overwrite each other.cssSelector(required): the selector string. With the defaultselectorType: cssit is a CSS selector, jQuery-style (#weight,.status,#status tr:nth-child(2) td:nth-child(2),[data-id]). WithselectorType: xpathit is an XPath expression (//table[@id='status']//tr[td='State']/td[2]). The name stayscssSelectorfor backwards compatibility with earlier versions of the node. Compiled at save; a malformed selector fails there, not at execution.selectorType(defaultcss): query language forcssSelector.cssis the jQuery-style default.xpathis worth reaching for when the value is picked by a text-content predicate (//tr[td='State']/td[2]), by axes (parent, ancestor), or against XML content where CSS is awkward.mode(defaulttext): how to extract from the matched element.text: the inner text of the first matched element, whitespace-trimmed.attr: the value of an attribute on the first matched element. Requiresattribute.html: the inner HTML of the first matched element, unmodified.
attribute: attribute name to read whenmodeisattr, for exampledata-value,href,class. Required whenmodeisattr; ignored otherwise.multiple(defaultfalse): whentrue, returns every match instead of the first. Withoutfields, the output is an array of extracted values. Withfields, an array of objects. Empty match always returns[](regardless ofemptyMatch).fields: only meaningful whenmultiple: true. Each entry describes one key on the per-match object; sub-selectors run within each outer match's DOM subtree. Nesting stops here — sub-fields have nomultipleorfieldsof their own.
Each fields entry mirrors the outer shape minus multiple/fields:
name(required): key on each per-match object; must be unique across siblings.cssSelector(required): selector string, run within the outer match. For XPath, use the relative descendant axis (.//td[1]) to keep the query scoped to the outer match; a bare//starts from the document root.selectorType(defaultcss):cssorxpath. Independent of the outer selector's type.mode(defaulttext):text,attr, orhtml.attribute: required whenmodeisattr.
emptyMatch (default null)
What to emit when a scalar selector finds no element (or, in attr mode, when the attribute is absent on the matched element):
null(default): keep the key on the output object with anullvalue, so the downstream shape stays stable.error: fail the node with the offending selector name in the error message.skip: omit the key from the output object entirely.
For multiple: true selectors, empty match always emits []; the policy above does not apply. Within multiple + fields, missing per-item sub-values follow the same policy.
Choose error only if the field is business-critical and the pipeline should stop on absence. Choose skip when the downstream node can gracefully handle missing keys.
maxInputBytes (default 1048576)
Maximum HTML input length in bytes. Inputs larger than this are refused at execution to keep a bad response from hoarding memory. Default is 1 MiB, which covers real device pages (a Zebra printer status page is about 10 KiB) with headroom. Raise it deliberately if the source serves a larger document.
Output
A flat object under $node["<name>"].result, keyed by each selector's name. Values are the extracted strings, or null when the selector matched nothing and emptyMatch is null.
Worked example: weight scale status page
The device serves this HTML at http://scale.local/status:
<div class="reading">
<span id="weight">42.5</span>
<span id="unit">kg</span>
</div>
<div class="status">stable</div>
Pipeline:
- REST GET
http://scale.local/status— the body ends up in$node["HTTP"].result.body. - HTML Extract with:
input={{ $node["HTTP"].result.body }}selectors=weight←#weight(mode text)unit←#unit(mode text)status←.status(mode text)
- Any downstream node reads
$node["HTML Extract"].result:
{
"weight": "42.5",
"unit": "kg",
"status": "stable"
}
Worked example: printer table status
Zebra-style printers often serve a table:
<table id="status">
<tr><td>State</td><td>Ready</td></tr>
<tr><td>Head</td><td>42°C</td></tr>
<tr><td>Labels</td><td>1247</td></tr>
</table>
Selectors dig into the cells:
state←#status tr:nth-child(1) td:nth-child(2)(text)headTemp←#status tr:nth-child(2) td:nth-child(2)(text)labelsLeft←#status tr:nth-child(3) td:nth-child(2)(text)
Output:
{
"state": "Ready",
"headTemp": "42°C",
"labelsLeft": "1247"
}
Worked example: XPath with a table-cell predicate
The Zebra table in the previous example is brittle to row reordering: #status tr:nth-child(1) pins the "State" row by position, so a firmware update that adds a "Firmware version" row at the top silently shifts every field. XPath's predicate matches on the label cell's text, so the extraction is stable:
<table id="status">
<tr><td>Firmware</td><td>2.7.1</td></tr>
<tr><td>State</td><td>Ready</td></tr>
<tr><td>Head</td><td>42°C</td></tr>
<tr><td>Labels</td><td>1247</td></tr>
</table>
Selectors (selectorType: xpath):
state←//table[@id='status']//tr[td='State']/td[2]headTemp←//table[@id='status']//tr[td='Head']/td[2]labelsLeft←//table[@id='status']//tr[td='Labels']/td[2]
Output:
{
"state": "Ready",
"headTemp": "42°C",
"labelsLeft": "1247"
}
The same expressions work regardless of the row order in the source HTML. Rows can appear, disappear or shuffle; as long as a row's label cell reads State, the neighbouring cell comes out on the state key.
Worked example: array of objects (alarm list)
Industrial dashboards routinely serve alarm lists, job queues and sensor grids as tables of items with per-row fields. Without array-of-objects support, extracting each row requires a fetch node, a split, a foreach and an aggregation — four hops that get the schema wrong easily. With multiple: true + fields:, one selector maps the whole shape.
Source HTML:
<table id="alarms">
<tr><td>10:32</td><td>HIGH</td><td>Temperature exceeded</td></tr>
<tr><td>10:35</td><td>MEDIUM</td><td>Low battery</td></tr>
<tr><td>10:41</td><td>HIGH</td><td>Sensor offline</td></tr>
</table>
One selector on the node:
name:alarmscssSelector:table#alarms trmultiple:truefields:timestamp←td:nth-child(1)(text)severity←td:nth-child(2)(text)message←td:nth-child(3)(text)
Output:
{
"alarms": [
{"timestamp": "10:32", "severity": "HIGH", "message": "Temperature exceeded"},
{"timestamp": "10:35", "severity": "MEDIUM", "message": "Low battery"},
{"timestamp": "10:41", "severity": "HIGH", "message": "Sensor offline"}
]
}
The outer selector picks each <tr>; each sub-selector runs within that row's DOM subtree, so sibling rows' cells never leak into another row's object. Downstream nodes iterate alarms directly with $node["HTML Extract"].result.alarms.
You can mix selector types: an outer xpath selector picking //table[@id='alarms']//tr with inner css fields for the cells works exactly the same way. For XPath in sub-fields, prefer the relative descendant axis (.//td[2]) so the query stays scoped to the outer match.
Worked example: sensor tile with data-* attributes
Modern device dashboards often carry values as attributes:
<div class="sensor" data-id="probe-3" data-value="23.4" data-unit="C">
Cold Storage Bay
</div>
Selectors use attr mode:
id←.sensor(attrdata-id)value←.sensor(attrdata-value)unit←.sensor(attrdata-unit)label←.sensor(text)
Output:
{
"id": "probe-3",
"value": "23.4",
"unit": "C",
"label": "Cold Storage Bay"
}
Notes
- CSS selectors use goquery (the Go equivalent of cheerio). XPath expressions use antchfx/htmlquery with the antchfx XPath 1.0 engine. Both parse against the same HTML tree (one call to
golang.org/x/net/html) so mixed selector types on the same node cost one parse. - The node does no framing above a single HTML document. If the source returns a stream of documents, reassemble upstream.
- Attribute values are returned exactly as stored in the DOM. Modern browsers decode HTML entities on display; if a specific device serves escaped attributes and you need the decoded value, decode downstream in a transform expression.
- For pages that require JavaScript execution to populate the value, this node returns whatever is in the initial HTML — usually an empty shell. Different problem, different tool.
- XPath in sub-fields: use the relative descendant axis (
.//tag) to keep the query scoped to each outer match. A bare//tagstarts from the document root and would pick up matches outside the outer scope.
Configuration reference
The fields below are generated from the node's config contract, so they match what the pipeline validator enforces and what the designer's form offers.
transform.extract.html
| Field | Type | Required | Default | Values | Description |
|---|---|---|---|---|---|
input | string | yes | — | accepts an expression | HTML input to parse. Resolves an expression like {{ $input[0].result.body }} or {{ $node["HTTP"].result.body }} to a string. Only expressions are accepted here; static text is not valid input for extraction. The pipeline is expected to fetch the HTML separately (REST, TCP, file) and hand its body to this node. |
selectors | object[] | yes | — | — | One entry per output field. Names must be unique; each entry names a key on the output object and a CSS selector that finds the value. |
emptyMatch | string | no | null | null, error, skip | What to emit when a selector matches nothing: null keeps the key with a null value (stable object shape); error fails the node with the offending selector; skip omits the key from the output. |
maxInputBytes | integer | no | 1.048576e+06 | at least 1 | Maximum HTML input length in bytes. Inputs larger than this are refused at execution to keep a bad response from hoarding memory. Default is 1 MiB (covers real device pages with headroom); raise it deliberately for larger documents. |