Skip to main content
Version: 3.0 (next)

AWS Glue AWS Glue Integration Guide

Connect to AWS Glue to run serverless ETL and read the metadata catalog the rest of your AWS analytics stack already queries. This guide covers connection setup, function configuration, and pipeline integration.

Overview​

AWS Glue is a serverless ETL service with a shared metadata catalog. The connector covers three things a pipeline needs from it:

  • Jobs — start an ETL job run with per-run arguments and worker overrides, then read its state, timings and error message
  • Crawlers — start a crawler so newly landed data becomes queryable, then read its state and the outcome of its last crawl
  • Data Catalog — browse databases, tables and full table schemas; the same catalog Athena, Redshift Spectrum and EMR read
  • Flexible authentication — IAM role / instance profile / IRSA via the AWS SDK default credential chain, or static access keys
  • Cross-account catalogs via an explicit Catalog ID
Started, not waited on

Glue jobs routinely run for tens of minutes and crawls for minutes. Start Job Run and Start Crawler return as soon as AWS accepts the request — they do not block until it finishes. Pair them with Get Job Run / Get Crawler in a later node to check the state, so a long run never holds a pipeline open.

Connection Configuration​

Creating an AWS Glue Connection​

Navigate to Connections → New Connection → AWS Glue and configure the following:

1. Profile Information​

FieldDefaultDescription
Profile Name-A descriptive name for this connection profile (required, max 100 characters)
Description-Optional description for this Glue connection

2. Region & Catalog​

FieldDefaultDescription
Regionus-east-1AWS region holding the Glue jobs, crawlers and Data Catalog (required). Glue is regional — all three live in one region
Catalog ID-AWS account ID owning the Data Catalog. Leave empty for the account the credentials belong to; set it to read a catalog shared from another account. Must be 12 digits

3. Authentication​

FieldDefaultDescription
Access Key ID-AWS Access Key ID. Masked on edit. Leave empty to use the AWS SDK default credential chain (env vars, shared config, IAM role, IRSA)
Secret Access Key-AWS Secret Access Key. Masked on edit; required when Access Key ID is set
Session Token-Session token for temporary STS credentials (optional). Masked on edit
Prefer IAM roles over static keys

Leave both key fields empty and the connector uses the AWS SDK default credential chain — an instance profile on EC2, IRSA on EKS, or ~/.aws/credentials locally. There is then no long-lived secret to rotate. Setting one key without the other is rejected at save time rather than silently falling back to the chain.

4. Advanced​

FieldDefaultDescription
Custom Endpoint-Custom Glue endpoint URL for moto, LocalStack or another Glue-compatible service. Leave empty for AWS Glue
Request Timeout30sTimeout for a single Glue API call (1s–1h). Bounds the API call, not the job run or crawl it starts
Max Retries3Retries after the first attempt for transient failures (0 = no retries, up to 10)

IAM Permissions​

The connection test calls glue:GetDatabases. Each function needs its own action:

FunctionIAM action
Start Job Runglue:StartJobRun
Get Job Runglue:GetJobRun
List Jobsglue:GetJobs
Start Crawlerglue:StartCrawler
Get Crawlerglue:GetCrawler
List Crawlersglue:GetCrawlers
List Databasesglue:GetDatabases
List Tablesglue:GetTables
Get Tableglue:GetTable

Testing the Connection​

Click Test Connection before saving. It calls GetDatabases for a single item — enough to prove the round trip, credentials and region. An empty catalog is still a successful test.

Failures name what to fix rather than echoing the SDK:

What you seeWhat it means
could not reach …Wrong endpoint URL, or the service is not running
… did not respond in timeThe endpoint answered too slowly: the call hit its timeout. Check the service's load, or raise the connection's Request Timeout or the function's Timeout
unable to resolve AWS credentialsNo keys given and the default chain found none
AWS rejected the provided credentialsThe Access Key ID or Secret Access Key is wrong
not authorized to call glue:GetDatabasesThe principal is valid but lacks the permission
Glue could not find the Data CatalogThe Catalog ID names an account whose catalog this principal cannot see

Function Builder​

Creating Glue Functions​

Open a saved connection, go to the Functions tab, and click New Function. The picker shows all nine operations, grouped by what they act on:

Select AWS Glue Function Type dialog

Choosing a Glue function type

Start Job Run Function​

Starts a run of an existing Glue job and returns the run ID.

FieldRequiredDescription
Job NameYesName of the Glue job to run. It must already exist in this region. Templatable
Job Arguments (JSON object)NoJSON object of job arguments, merged over the job's defaults. Glue expects the -- prefix on each key, and every value must be a string. Templatable
Worker TypeNoOverride the job's worker type for this run (G.025X, G.1X, G.2X, G.4X, G.8X, Z.2X, Standard)
Number of WorkersNoOverride the number of workers for this run (2–1000)
TimeoutNoBound on the StartJobRun API call itself

Result:

{
"runId": "jr_8a1f3c5e9b2d4a6f8c0e2a4b6d8f0a2c4e6b8d0f2a4c6e8b0d2f4a6c8e0b2d4f"
}
Worker overrides come as a pair

Glue takes WorkerType and NumberOfWorkers together — one without the other is an InvalidInputException at run time. Set both, or leave both empty to use the job's own settings. The form and the API both refuse a half-filled pair at save time.

Job arguments are strings

{"--retries": 3} is rejected: Glue's job arguments are a string→string map, so write {"--retries": "3"}. The error names the offending key.

Get Job Run Function​

Reads one job run by job name and run ID. This is the polling half of Start Job Run.

FieldRequiredDescription
Job NameYesName of the Glue job the run belongs to. Templatable
Run IDYesThe run ID returned by Start Job Run. Templatable
TimeoutNoBound on this single operation

Result:

{
"runId": "jr_8a1f…",
"state": "SUCCEEDED",
"startedOn": "2026-09-22T02:00:00Z",
"completedOn": "2026-09-22T02:07:13Z",
"executionTime": 433,
"attempt": 0,
"errorMessage": "",
"workerType": "G.1X",
"numberOfWorkers": 2,
"glueVersion": "4.0"
}

errorMessage is always present and empty on a run that has not failed. completedOn is empty while the run is still going — that is how a polling loop tells a terminal state from a running one without parsing the state string.

List Jobs Function​

Returns the Glue jobs defined in the region with their role, command, Glue version and worker shape.

FieldRequiredDescription
Max ItemsNoMaximum jobs to return, 1–1000 (default 100). The connector paginates in pages of 100 until the budget is met
TimeoutNoBound on this single operation

Result:

{
"jobs": [
{
"name": "ot-daily-rollup",
"description": "Daily OT rollup",
"role": "arn:aws:iam::123456789012:role/GlueETL",
"glueVersion": "4.0",
"workerType": "G.1X",
"numberOfWorkers": 2,
"commandName": "glueetl",
"scriptLocation": "s3://ot-scripts/rollup.py",
"createdOn": "2026-04-02T09:12:00Z"
}
],
"count": 1,
"nextToken": "AAAA…"
}

nextToken appears only when the item budget stopped the listing; _metadata.truncated says the same thing as a boolean.

Start Crawler Function​

Starts an existing crawler so it re-reads its targets and folds any new tables or partitions into the catalog.

FieldRequiredDescription
Crawler NameYesName of the crawler to start. It must already exist in this region. Templatable
TimeoutNoBound on the StartCrawler API call itself

Result:

{ "started": true }
Already running is transient, not fatal

Glue refuses to start a crawler that is already running, and the node reports that as a failure — but a transient one, so the run is retried rather than dead-lettered. Two pipelines refreshing the same catalog hit this routinely. Use Get Crawler to wait for READY before starting again.

Get Crawler Function​

Reads one crawler by name, with the outcome of its last crawl.

FieldRequiredDescription
Crawler NameYesName of the crawler to read. Templatable
TimeoutNoBound on this single operation

Result:

{
"state": "READY",
"databaseName": "ot_archive",
"description": "Discovers new OT partitions",
"role": "arn:aws:iam::123456789012:role/GlueCrawler",
"crawlElapsedTime": 0,
"lastCrawlStatus": "SUCCEEDED",
"lastCrawlStartedOn": "2026-09-22T01:30:00Z",
"lastCrawlErrorMessage": ""
}

The three lastCrawl* keys are always present and empty on a crawler that has never run.

List Crawlers Function​

Returns the crawlers in the region with their state, target database and last-crawl status.

FieldRequiredDescription
Max ItemsNoMaximum crawlers to return, 1–1000 (default 100)
TimeoutNoBound on this single operation

Result:

{
"crawlers": [
{
"name": "ot-archive-crawler",
"state": "READY",
"databaseName": "ot_archive",
"description": "Discovers new OT partitions",
"lastCrawlStatus": "SUCCEEDED",
"lastCrawlStartedOn": "2026-09-22T01:30:00Z"
}
],
"count": 1,
"nextToken": "AAAA…"
}

List Databases Function​

Returns the databases in the Data Catalog.

FieldRequiredDescription
Max ItemsNoMaximum databases to return, 1–1000 (default 100)
TimeoutNoBound on this single operation

Result:

{
"databases": [
{
"name": "ot_archive",
"description": "Landed OT telemetry",
"locationUri": "s3://ot-lake/archive/",
"catalogId": "123456789012"
}
],
"count": 1,
"nextToken": "AAAA…"
}

List Tables Function​

Returns the tables in one catalog database, optionally filtered by a name pattern.

FieldRequiredDescription
DatabaseYesCatalog database whose tables to list. Templatable
Name FilterNoGlue filter pattern on the table name, e.g. readings_*. Leave empty for all. Templatable
Max ItemsNoMaximum tables to return, 1–1000 (default 100)
TimeoutNoBound on this single operation

Result:

{
"tables": [
{
"name": "readings",
"tableType": "EXTERNAL_TABLE",
"location": "s3://ot-lake/archive/readings/",
"columnCount": 4,
"partitionKeyCount": 1,
"updateTime": "2026-09-22T01:34:00Z"
}
],
"count": 1,
"nextToken": "AAAA…"
}

A listing carries counts rather than the columns themselves — a wide catalog would otherwise deliver thousands of column objects nobody asked for. Get Table returns the schema.

Get Table Function​

Reads a single table's full schema.

FieldRequiredDescription
DatabaseYesCatalog database the table belongs to. Templatable
TableYesName of the table to read. Templatable
TimeoutNoBound on this single operation

Result:

{
"tableType": "EXTERNAL_TABLE",
"location": "s3://ot-lake/archive/readings/",
"inputFormat": "org.apache.hadoop.mapred.TextInputFormat",
"outputFormat": "org.apache.hadoop.hive.ql.io.HiveIgnoreKeyTextOutputFormat",
"columns": [
{ "name": "machine_id", "type": "string", "comment": "Line asset ID" }
],
"partitionKeys": [
{ "name": "day", "type": "string", "comment": "" }
],
"createTime": "2026-04-02T09:12:00Z",
"updateTime": "2026-09-22T01:34:00Z"
}

columns and partitionKeys are always lists — empty rather than absent, so a downstream ForEach over an unpartitioned table does nothing instead of reporting source is nil.

Using Parameters​

Any field marked templatable accepts ((parameterName)). Parameters are detected as you type and listed in the Function Parameters card:

Auto-detected Glue function parameters

Parameters detected from ((day)) and ((mode)) in the job arguments

A pipeline supplies their values at execution time, so one Start Job Run function can serve every day of a backfill.

Pipeline Integration​

Each function becomes a node in the pipeline editor under Databases. See the AWS Glue node reference for node-level configuration.

The shape that comes up most often is start → wait → read:

  1. Start Job Run kicks off the ETL job and emits runId
  2. A Delay node waits
  3. Get Job Run reads the state using $node["Start Job Run"].result.runId
  4. A Condition node branches on state — loop back to the delay while it is RUNNING, continue on SUCCEEDED, alert on FAILED

Common Use Cases​

Refreshing the catalog after landing files​

A pipeline writes new partitions to S3, then Start Crawler makes them queryable. A later Get Crawler confirms the crawl succeeded before an Athena query runs against the new tables.

Nightly rollup with per-day arguments​

Start Job Run with {"--day": "((day))"} runs the same job for whichever day the trigger supplies. The run ID goes into an execution log; Get Job Run reads the outcome on the next pass.

Schema-driven mapping​

Get Table reads a table's columns from the catalog, and the pipeline maps them onto UNS topics — so a column added upstream flows through without a pipeline edit.

Troubleshooting​

SymptomCauseFix
Glue found no job named "X" in region …Wrong job name, or the job lives in another regionUse List Jobs on the same connection to see what exists
Glue found no crawler named "X"Wrong crawler name or regionUse List Crawlers
Glue found no table "X" in database "Y"Wrong name, region, or Catalog IDUse List Databases and List Tables; check the Catalog ID on the connection
crawler "X" is already runningAnother pipeline started itExpected; the run is retried. Gate on Get Crawler reporting READY
job "X" is already at its maximum concurrent runsThe job's concurrency limit is reachedWait for a run to finish, or raise the limit on the job
workerType and numberOfWorkers must be set togetherOnly one of the worker overrides is filledSet both, or clear both
value of "--x" is a numberA job argument is not a stringQuote it: {"--x": "3"}
Listing looks shortThe Max Items budget stopped itRaise Max Items; _metadata.truncated and nextToken say where it stopped