Skip to main content
Version: 3.0 (next)

AWS Glue Nodes

MaestroHub provides native AWS Glue integration for serverless ETL and metadata discovery. Use these nodes to start Glue jobs and crawlers from a pipeline, follow how they went, and read the Data Catalog that Athena, Redshift Spectrum and EMR already query.

Configuration Quick Reference​

FieldWhat you chooseDetails
ParametersConnection, Function, Function Parameters, Timeout OverrideSelect the connection profile, function, configure function parameters with expression support, and optionally override the timeout.
SettingsDescription, Timeout (seconds), Retry on Timeout, Retry on Fail, On ErrorNode description, maximum execution time, retry behavior on timeout or failure, and error handling strategy. All execution settings default to pipeline-level values.

Node Types​

The Glue connector provides nine node types across jobs, crawlers and the Data Catalog:

NodePurposeCommon Use Cases
Start Job RunStart an ETL job run and return its run IDNightly rollups, per-day backfills, catch-up loads on bigger workers
Get Job RunRead a run's state, timings and error messageWaiting for a job, branching on the outcome, alerting on failure
List JobsList the job definitions in the regionDiscovery before wiring a run, ETL inventory
Start CrawlerStart a crawler so new data enters the catalogRefreshing the catalog after landing files, picking up a new partition
Get CrawlerRead a crawler's state and last crawlWaiting for READY, surfacing a failed crawl
List CrawlersList the crawlers in the regionDiscovery, finding crawlers whose last run failed
List DatabasesList the Data Catalog's databasesDiscovery, iterating over databases
List TablesList a database's tables, optionally filteredCatalog-driven iteration, confirming a crawler's output
Get TableRead one table's full schemaSchema-driven mapping, detecting a schema change
Started, not waited on

Start Job Run and Start Crawler return as soon as AWS accepts the request. Glue jobs routinely run for tens of minutes, so a node that blocked until completion would hold the whole pipeline open. Pair each Start node with its Get node in a later step.


AWS Glue Start Job Run node

AWS Glue Start Job Run Node

AWS Glue Start Job Run Node​

Starts a run of an existing Glue job with optional per-run arguments and worker overrides, and returns the run ID immediately.

Configuration: Job Name (required, templatable), Job Arguments (JSON object of string values, templatable), Worker Type and Number of Workers (set together or not at all).

Output

KeyWhat it is
result.runIdThe run ID Glue assigned — pass it to Get Job Run to follow the run

A later node reads it as $node["Start Job Run"].result.runId.


AWS Glue Get Job Run node

AWS Glue Get Job Run Node

AWS Glue Get Job Run Node​

Reads one job run by job name and run ID. This is the polling half of Start Job Run — loop over it until the state is terminal.

Configuration: Job Name and Run ID (both required, both templatable).

Output

KeyWhat it is
result.runIdThe run this is about
result.stateRUNNING, SUCCEEDED, FAILED, TIMEOUT, STOPPED or another Glue run state
result.startedOnWhen the run started, RFC3339
result.completedOnWhen it finished, RFC3339 — empty while it is still running
result.executionTimeSeconds of billed execution
result.attemptWhich retry attempt this run is, 0 for the first
result.errorMessageWhy it failed — empty on a run that has not failed
result.workerTypeThe worker size the run used
result.numberOfWorkersHow many workers the run used
result.glueVersionThe Glue version the run used

result.completedOn being empty is how a loop tells a running state from a terminal one without parsing result.state.


AWS Glue List Jobs node

AWS Glue List Jobs Node

AWS Glue List Jobs Node​

Lists the Glue job definitions in the configured region.

Configuration: Max Items (1–1000, default 100).

Output

KeyWhat it is
result.jobsOne object per job — name, description, role, glueVersion, workerType, numberOfWorkers, commandName, scriptLocation and createdOn
result.countHow many were listed
result.nextTokenWhere the listing stopped — present only when _metadata.truncated is true

AWS Glue Start Crawler node

AWS Glue Start Crawler Node

AWS Glue Start Crawler Node​

Starts an existing crawler so it re-reads its targets and folds any new tables or partitions into the Data Catalog.

Configuration: Crawler Name (required, templatable).

Output

KeyWhat it is
result.startedGlue accepted the start — Get Crawler is what reports how the crawl went
Already running is retried, not dead-lettered

Glue refuses to start a crawler that is already running. The node fails, but the failure is classified transient, so the pipeline's retry policy applies. Two pipelines refreshing the same catalog hit this routinely.


AWS Glue Get Crawler node

AWS Glue Get Crawler Node

AWS Glue Get Crawler Node​

Reads one crawler by name, with the outcome of its last crawl.

Configuration: Crawler Name (required, templatable).

Output

KeyWhat it is
result.stateREADY, RUNNING or STOPPING
result.databaseNameThe catalog database the crawler writes into
result.descriptionThe crawler's description
result.roleThe IAM role the crawler runs as
result.crawlElapsedTimeMilliseconds the running crawl has taken so far, 0 when it is not running
result.lastCrawlStatusSUCCEEDED, CANCELLED or FAILED — empty on a crawler that has never run
result.lastCrawlStartedOnWhen the last crawl started, RFC3339 — empty on a crawler that has never run
result.lastCrawlErrorMessageWhy the last crawl failed — empty when it did not

AWS Glue List Crawlers node

AWS Glue List Crawlers Node

AWS Glue List Crawlers Node​

Lists the crawlers in the configured region.

Configuration: Max Items (1–1000, default 100).

Output

KeyWhat it is
result.crawlersOne object per crawler — name, state, databaseName, description, lastCrawlStatus and lastCrawlStartedOn
result.countHow many were listed
result.nextTokenWhere the listing stopped — present only when _metadata.truncated is true

AWS Glue List Databases node

AWS Glue List Databases Node

AWS Glue List Databases Node​

Lists the databases in the Glue Data Catalog — the starting point for discovering what is queryable across the AWS analytics stack.

Configuration: Max Items (1–1000, default 100).

Output

KeyWhat it is
result.databasesOne object per database — name, description, locationUri and catalogId
result.countHow many were listed
result.nextTokenWhere the listing stopped — present only when _metadata.truncated is true

AWS Glue List Tables node

AWS Glue List Tables Node

AWS Glue List Tables Node​

Lists the tables in one catalog database, optionally filtered by a Glue name pattern.

Configuration: Database (required, templatable), Name Filter (e.g. readings_*, templatable), Max Items (1–1000, default 100).

Output

KeyWhat it is
result.tablesOne object per table — name, tableType, location, columnCount, partitionKeyCount and updateTime. A listing carries counts rather than the columns themselves; Get Table returns the schema
result.countHow many were listed
result.nextTokenWhere the listing stopped — present only when _metadata.truncated is true

AWS Glue Get Table node

AWS Glue Get Table Node

AWS Glue Get Table Node​

Reads a single table from the Data Catalog with its full schema.

Configuration: Database and Table (both required, both templatable).

Output

KeyWhat it is
result.tableTypeEXTERNAL_TABLE, VIRTUAL_VIEW or another Hive table type
result.locationWhere the data sits
result.inputFormatThe Hive input format
result.outputFormatThe Hive output format
result.columnsOne object per column — name, type and comment
result.partitionKeysOne object per partition key — name, type and comment. Empty on an unpartitioned table
result.createTimeWhen the table was created, RFC3339
result.updateTimeWhen it was last updated, RFC3339

result.columns and result.partitionKeys are always lists, empty rather than absent — a ForEach over an unpartitioned table does nothing instead of reporting source is nil.


Node Metadata​

Every Glue node's _metadata carries the connector's own facts about the call:

KeyWhat it is
_metadata.methodThe operation that ran, e.g. glue.start_job_run
_metadata.connectionIdThe connection profile that ran it
_metadata.protocolAlways glue
_metadata.jobNameThe job this call addressed (job operations)
_metadata.crawlerNameThe crawler this call addressed (crawler operations)
_metadata.databaseNameThe catalog database this call addressed (table operations)
_metadata.tableNameThe table this call addressed (Get Table)
_metadata.truncatedWhether the item budget stopped the listing (listing operations)

Pipeline Patterns​

Start, wait, read​

The shape that comes up most often:

  1. Start Job Run emits result.runId
  2. A Delay node waits
  3. Get Job Run reads the state with $node["Start Job Run"].result.runId
  4. A Condition branches on result.state — loop back while RUNNING, continue on SUCCEEDED, alert on FAILED

Land, crawl, query​

A pipeline writes new partitions to S3, Start Crawler makes them queryable, Get Crawler confirms result.lastCrawlStatus is SUCCEEDED, and an Athena query runs against the fresh tables.

Catalog-driven iteration​

List Tables feeds a ForEach, and Get Table reads each table's result.columns so a mapping follows the catalog rather than a hard-coded column list.