List Files
Available Since
- v5.2.38 and later
The List Files task is used to retrieve a list of files from a specific location. The task determines the source type from the URL scheme and lists all files at that location. It supports AWS S3 and Google Cloud Storage buckets.
During execution, the task detects the input type from the URL scheme. For example, a URL starting with s3:// is treated as an AWS S3 bucket, and a URL starting with gs:// is treated as a Google Cloud Storage bucket. The task lists files (filtered by file type if provided) and returns the absolute paths as an array. If you specify an output location, the task also writes the list to that location.
Prerequisites
If the location of the file is not publicly available, you must create an appropriate integration with the required access keys or tokens. Integrate the following with Orkes Conductor, depending on your source:
Task parameters
Configure these parameters for the List Files task.
| Parameter | Description | Required/ Optional |
|---|---|---|
| inputParameters.inputLocation | The location of the files to be listed. Example based on the integration type:
|
Required. |
| inputParameters.integrationName | If the location of the file to be listed is not publicly available, select the integration name of the Cloud Providers integrated with your Conductor cluster. Note: If you haven’t configured any integration on your Orkes Conductor cluster, go to the Integrations tab and configure the required Cloud Providers. |
Optional. |
| inputParameters.fileTypes | The file types to be listed. If omitted, all file types are included. Supported values:
|
Optional. |
| inputParameters.outputLocation | The storage location where the resulting file list is saved as a text file, with each line in the text file containing one absolute file path. For example:
Note: Cloud storage output requires a corresponding integration with write permissions. Use integrationNames parameter (e.g., {"aws": "my-integration"}) to add any additional integration configurations. |
Optional. |
| inputParameters.integrationNames | A key-value map of integration types and names. Use this when multiple integrations are needed. The key represents the type of integration (for example, aws, gcp), and the value specifies the name of the corresponding integration. |
Optional. |
The following are generic configuration parameters that can be applied to the task and are not specific to the List Files task.
Other generic parameters
Here are other parameters for configuring the task behavior.
| Parameter | Description | Required/ Optional |
|---|---|---|
| optional | Whether the task is optional. If set to true, any task failure is ignored, and the workflow continues with the task status updated to COMPLETED_WITH_ERRORS. However, the task must reach a terminal state. If the task remains incomplete, the workflow waits until it reaches a terminal state before proceeding. |
Optional. |
Task configuration
This is the task configuration for a List Files task.
{
"name": "list_files",
"taskReferenceName": "lf",
"type": "LIST_FILES",
"inputParameters": {
"inputLocation": "<YOUR-LOCATION>",
"fileTypes": ["<YOUR-FILE-TYPE>"]
}
}
Task output
The List Files task will return the following parameters.
| Parameter | Description |
|---|---|
| files | An array of absolute file paths listed from the input location. |
Examples
Here are some examples for using the List Files task.
List files from an AWS S3 bucket
In this example, we will:
- Create an AWS integration in Orkes Conductor.
- Create a workflow that uses the List Files task.
- Run the workflow and verify the output.
Step 1: Create an AWS integration in Orkes Conductor
Create an AWS integration with the connection type as Access Key/Secret. Ensure that the access key has read access to the S3 bucket from which the files are to be listed.
Step 2: Create a workflow in Orkes Conductor
To create a workflow definition using Conductor UI:
- Go to Definitions > Workflow, from the left navigation menu on your Conductor cluster.
- Select + Define workflow.
- In the Code tab, paste the following code:
Workflow definition:
{
"name": "list_files_s3",
"description": "Workflow to list PDF files from an AWS S3 bucket.",
"version": 1,
"tasks": [
{
"name": "list_files",
"taskReferenceName": "list_files_ref",
"inputParameters": {
"inputLocation": "<YOUR-S3-FOLDER-URI>",
"fileTypes": [
"pdf"
],
"integrationName": "<YOUR-AWS-INTEGRATION-NAME>"
},
"type": "LIST_FILES"
}
],
"schemaVersion": 2
}
- Replace
<YOUR-S3-FOLDER-URI>with the S3 URI of the folder, for examples3://<YOUR-BUCKET-NAME>/<FOLDER-NAME>/, and<YOUR-AWS-INTEGRATION-NAME>with the integration name created in Step 1. - Select Save > Confirm.
Step 3: Run the workflow and verify the output
Select the Execute button from the workflow definition page. Once the workflow is completed, the task output contains a files array with the S3 paths of all PDF files in the folder:
Each element in the files array represents one file path from the input location. These can be used directly by downstream tasks for further processing.
List files from an AWS S3 bucket and save the list to S3
In this example, we will:
- Create an AWS integration in Orkes Conductor.
- Create a workflow that uses the List Files task.
- Run the workflow and verify the output.
Step 1: Create an AWS integration in Orkes Conductor
Create an AWS integration with the connection type as Access Key/Secret. Ensure that the access key has read access to the S3 bucket from which the files are to be listed, and write access to the S3 bucket where the output text file will be stored.
Step 2: Create a workflow in Orkes Conductor
To create a workflow using Conductor UI:
- Go to Definitions > Workflow from the left navigation menu on your Conductor cluster.
- Select + Define workflow.
- In the Code tab on the right panel, paste the following code:
{
"name": "list_files_s3_save",
"description": "Workflow to list PDF files from an AWS S3 bucket and save the list to a text file in S3.",
"version": 1,
"tasks": [
{
"name": "list_files",
"taskReferenceName": "list_files_ref",
"inputParameters": {
"inputLocation": "<YOUR-S3-FOLDER-URI>",
"fileTypes": [
"pdf"
],
"integrationName": "<YOUR-AWS-INTEGRATION-NAME>",
"outputLocation": "<YOUR-S3-TXT-FILE-URI>"
},
"type": "LIST_FILES"
}
],
"schemaVersion": 2
}
- Replace the following:
<YOUR-S3-FOLDER-URI>with the S3 URI of the folder from which files are to be listed. For example,s3://<YOUR-BUCKET-NAME>/<FOLDER-NAME>/.<YOUR-AWS-INTEGRATION-NAME>with the integration name created in Step 1.<YOUR-S3-TXT-FILE-URI>with the S3 URI of the text file where the list will be saved. For example,s3://<YOUR-BUCKET-NAME>/file-list.txt.
- Select Save > Confirm.
Step 3: Run the workflow and verify the output
Select the Execute button from the workflow definition page. Once the workflow is completed, the task output contains a files array with the S3 paths of all PDF files in the folder:
The list is also written to the text file at the output location. Open the S3 bucket to verify that the text file has been created, with each line containing one absolute file path.