Skip to main content
Version: 3.0 (next)
Convert to File node

Convert to File node

Convert to File Node

Overview​

  • Type: transform.data.serializer
  • Display Name: Convert to File
  • Category: transform
  • Execution: supportsExecution: true
  • I/O Handles:
    • Input: In (left)
    • Output: Out (right)

Purpose: Convert structured data (arrays of objects or arrays) into CSV, Excel, or Parquet file format. Supports header management, data formatting options, append mode for streaming scenarios, and base64 encoding for safe transport to file writers or connectors.


Format Configuration​

ParameterTypeDefaultRequiredDescription
Formatselect"csv"YesOutput format: csv, excel, or parquet.

CSV-Specific​

ParameterTypeDefaultRequiredDescription
Delimiterstring","Yes (CSV)Field separator character. Common values: , ; \t `
Line Endingselect"LF"NoLine ending style: LF (Unix/Linux) or CRLF (Windows).

Excel-Specific​

ParameterTypeDefaultRequiredDescription
Sheet Namestring"Sheet1"Yes (Excel)Worksheet name. Max 31 characters. Cannot contain: / \ ? * [ ].
Auto FilterbooleantrueNoAdd dropdown filter to header row.
Freeze HeaderbooleantrueNoFreeze the header row for scrolling (freeze panes).

Parquet-Specific​

ParameterTypeDefaultRequiredDescription
Parquet Column Schemalist of {name, type}derived from headersNoDeclares the output column types. type is one of string, int64, float64, bool, timestamp. When omitted, every column is written as string.

Notes:

  • Declare the schema in production. Without it every column is a string, which costs you predicate pushdown and correct aggregation in whatever queries the file — the main reason to choose Parquet over CSV in the first place.
  • Cells that fail their declared type's parser become NULL rather than failing the write. The count is reported as metadata.parseFailures, and metadata.parseFailureRate gives it as a fraction of typed cells written this execution. A rate that climbs is upstream schema drift.
  • timestamp cells are parsed as RFC 3339 (with or without fractional seconds) and stored as UTC microseconds — the highest precision Spark and Delta Lake can read. Sub-microsecond digits are truncated.
  • Empty cells become NULL and do not count as parse failures — they are explicit absences.

Two output metadata fields are added for Parquet only:

FieldDescription
_metadata.parseFailuresCount of cells that became NULL because they failed the declared column's parser.
_metadata.parseFailureRateparseFailures over typed cells written this execution. 0.0 when there are no typed columns.
Parquet output is always base64

Parquet is binary, so the node base64-encodes it regardless of the Encoding setting. Any downstream writer must be told to decode it: the OneLake write function needs Input Format = base64, and the local file and SMB writers need the same. Left as text, the base64 characters are written as the file contents and the result is not a Parquet file.

Header Configuration​

ParameterTypeDefaultRequiredDescription
Header Modeselect"auto"YesHow headers are determined (see below).
Headersstring[][]If providedCustom column header names.
Include HeadersbooleantrueNoWrite header row to output file.
Always Include HeadersbooleanfalseNoForce headers on every execution (ignores state).
Strict HeadersbooleanfalseNoValidate data structure matches established headers.

Header Mode Options​

ModeDescription
autoAutomatically detect headers from input data structure (object keys).
providedUse the custom headers array configured in the node.
fromFirstRowTreat first row/element of input as headers (for array of arrays).
noneNo headers — output data rows only.

Input / Output Configuration​

ParameterTypeDefaultRequiredDescription
Input Formatselect"auto"NoExpected input structure: auto, arrayOfArrays, or arrayOfObjects.
First Row Is HeaderbooleanfalseNoTreat first array element as headers (for array of arrays input).
Input Fieldstring"result"YesWhere to read the rows from: a path such as result or result.rows, or an expression such as {{ $node["Read"].result.rows }} to reach any upstream node.
Encodingselect"base64"NoOutput encoding: base64 or utf8. Excel and Parquet always use base64 regardless of this setting — they are binary.
Append ModebooleantrueNoAccumulate data across multiple executions.

Data Processing​

ParameterTypeDefaultDescription
Null Valuestring""String representation for null/undefined values.
Bool Formatselect"true/false"Boolean formatting: true/false, TRUE/FALSE, or 1/0.

Settings​

SettingOptionsDefaultDescription
Timeout (seconds)numberPipeline defaultMaximum execution time for this node (1–600).
Retry on TimeoutPipeline Default / Enabled / DisabledPipeline DefaultWhether to retry on timeout.
Retry on FailPipeline Default / Enabled / DisabledPipeline DefaultWhether to retry on failure. When Enabled, shows Advanced Retry Configuration.
On ErrorPipeline Default / Stop Pipeline / Continue ExecutionPipeline DefaultBehavior when node fails after all retries.

Advanced Retry Configuration​

Only visible when Retry on Fail is set to Enabled.

FieldTypeDefaultRangeDescription
Max Attemptsnumber31–10Maximum retry attempts.
Initial Delay (ms)number1000100–30,000Wait before first retry.
Max Delay (ms)number1200001,000–300,000Upper bound for backoff delay.
Multipliernumber2.01.0–5.0Exponential backoff multiplier.
Jitter Factornumber0.10–0.5Random jitter.

Output Format​

Success Packet​

{
"result": "SGVsbG8sV29ybGQKSm9obiwzMApKYW5lLDI1Cg==",
"_metadata": {
"format": "csv",
"rowsProcessed": 2,
"totalRows": 2,
"headersIncluded": true,
"headers": ["Name", "Age"],
"encoding": "base64"
}
}
FieldDescription
resultEncoded file content. The file is always in result; read it downstream as {{ $node["Convert to File"].result }} (use your node's name).
_metadata.formatOutput format (csv, excel, or parquet).
_metadata.rowsProcessedNumber of rows processed in this execution.
_metadata.totalRowsCumulative row count across all executions (append mode).
_metadata.headersIncludedWhether headers were written in this execution.
_metadata.headersColumn headers used.
_metadata.encodingEncoding used (base64 or utf8).

Errors​

When the node fails, the error message is shown on the node in the execution history. Common errors:

  • Configuration error (invalid format, missing required fields)
  • Input field not found
  • Header validation failed (strict mode)
  • Failed to generate CSV/Excel/Parquet

Validation Rules​

  • label: required (non-empty)
  • format: required; must be csv, excel, or parquet
  • CSV: delimiter required and non-empty
  • Excel: sheetName required; must not contain / \ ? * [ ]; max 31 characters
  • headerMode: must be one of auto | provided | fromFirstRow | none
  • When headerMode === 'provided': headers must be a non-empty array
  • inputField: required (non-empty)

Manual Actions​

This node supports manual actions accessible from the node context menu:

ActionDescription
ResetClear accumulated data and reset state. Use when switching output files or starting fresh.
Get StateView current serializer state including row counts, detected headers, and accumulated data.


Configuration reference​

The fields below are generated from the node's config contract, so they match what the pipeline validator enforces and what the designer's form offers.

transform.data.serializer​

FieldTypeRequiredDefaultValuesDescription
formatstringnocsvcsv, excel, parquetFile format to produce: csv (text), excel (.xlsx workbook) or parquet (columnar, typed)
delimiterstringno,—csv only: the field separator, exactly one character, e.g. "," ";" "|" or a real tab character
lineEndingstringnoLFLF, CRLFcsv only: LF ends lines with \n (Unix), CRLF with \r\n (Windows)
headerModestringnoautoauto, provided, fromFirstRow, noneWhere column names come from: auto takes the object keys of the input rows; provided uses the headers list; fromFirstRow treats the first input row as the names (array-of-arrays input); none writes no header row
headersstring[]no——Column names, in order, used when headerMode is provided; must be unique. Ignored in the other modes
includeHeadersbooleannotrue—Write a header row; with appendMode the row is written once, on the first execution
alwaysIncludeHeadersbooleannofalse—Write the header row on every execution, ignoring whether an earlier execution already wrote it
strictHeadersbooleannofalse—Fail an execution whose columns differ from the first execution's; for parquet with parquetColumns, also fail when a declared column is missing from the input instead of writing null
inputFormatstringnoautoauto, arrayOfArrays, arrayOfObjects, primitiveArrayShape of the rows at inputField: auto detects from the first element; arrayOfArrays is rows of cells; arrayOfObjects is rows keyed by column name; primitiveArray is one column of scalars
firstRowIsHeaderbooleannofalse—arrayOfArrays input only: use the first row as column names instead of data (same effect as headerMode fromFirstRow)
inputFieldstringnoresultaccepts an expressionDot path to the rows inside the node input, or an expression. Every node output is {result, _metadata}, so paths start at result, e.g. result or result.rows; an expression such as {{ $node["Read"].result.rows }} reaches any upstream node, which a plain path cannot; the execution fails when the path is missing. When the input is not an object (an array from upstream, or several predecessors) the input itself is used and the path is ignored
nullValuestringno——Text written for null cells and for the missing cells of short rows; empty by default
boolFormatstringnotrue/falsetrue/false, TRUE/FALSE, 1/0How boolean cells are written
flattenNestedbooleannofalse—Accepted but has no effect: nested objects and arrays are always written as JSON text in their cell
sheetNamestringnoSheet1—excel only: worksheet name, at most 31 characters and none of / \ ? * [ ]
autoFilterbooleannotrue—excel only: add a filter dropdown to the header row
freezeHeaderbooleannotrue—excel only: freeze the header row so it stays visible while scrolling
parquetColumnsobject[]no——parquet only: the typed column schema in output order; names must be unique and each must match an input column name (a missing one is null). Omit it to write every input column as a string
outputFieldstringno——Accepted but ignored: the file content is always the node's result ($node["Name"].result). Kept so saved pipelines keep loading
encodingstringnobase64base64, utf8How the file bytes are emitted: base64 text, or utf8 plain text. excel and parquet are always base64 whatever is set here
appendModebooleannotrue—Accumulate rows across executions and emit the whole file each time; false emits only this execution's rows. The Reset action clears the accumulated rows