> For the complete documentation index, see [llms.txt](https://docs.selfuel.digital/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.selfuel.digital/data-integration-with-nexus/nexus-elements/connectors/source/s3file.md).

# S3File

> S3 File Source Connector

### Key Features[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#key-features) <a href="#key-features" id="key-features"></a>

* [x] &#x20;batch
* [ ] &#x20;stream
* [x] &#x20;exactly-once

Read all the data in a split in a pollNext call. What splits are read will be saved in snapshot.

* [x] &#x20;column projection
* [x] &#x20;parallelism
* [ ] &#x20;support user-defined split
* [x] &#x20;file format type
  * [x] &#x20;text
  * [x] &#x20;csv
  * [x] &#x20;parquet
  * [x] &#x20;orc
  * [x] &#x20;json
  * [x] &#x20;excel
  * [x] &#x20;xml
  * [x] &#x20;binary

### Description[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#description) <a href="#description" id="description"></a>

Read data from aws s3 file system.

### Supported DataSource Info[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#supported-datasource-info) <a href="#supported-datasource-info" id="supported-datasource-info"></a>

| Datasource | Supported versions |
| ---------- | ------------------ |
| S3         | current            |

### Data Type Mapping[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#data-type-mapping) <a href="#data-type-mapping" id="data-type-mapping"></a>

Data type mapping is related to the type of file being read, We supported as the following file types:

`text` `csv` `parquet` `orc` `json` `excel` `xml`

#### JSON File Type[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#json-file-type) <a href="#json-file-type" id="json-file-type"></a>

If you assign file type to `json`, you should also assign schema option to tell connector how to parse data to the row you want.

For example:

upstream data is the following:

```

{"code":  200, "data":  "get success", "success":  true}

```

You can also save multiple pieces of data in one file and split them by newline:

```

{"code":  200, "data":  "get success", "success":  true}
{"code":  300, "data":  "get failed", "success":  false}

```

you should assign schema as the following:

```

schema {
    fields {
        code = int
        data = string
        success = boolean
    }
}

```

connector will generate data as the following:

| code | data        | success |
| ---- | ----------- | ------- |
| 200  | get success | true    |

#### Text Or CSV File Type[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#text-or-csv-file-type) <a href="#text-or-csv-file-type" id="text-or-csv-file-type"></a>

If you assign file type to `text` `csv`, you can choose to specify the schema information or not.

For example, upstream data is the following:

```

tyrantlucifer#26#male

```

If you do not assign data schema connector will treat the upstream data as the following:

| content               |
| --------------------- |
| tyrantlucifer#26#male |

If you assign data schema, you should also assign the option `field_delimiter` too except CSV file type

you should assign schema and delimiter as the following:

```

field_delimiter = "#"
schema {
    fields {
        name = string
        age = int
        gender = string 
    }
}

```

connector will generate data as the following:

| name          | age | gender |
| ------------- | --- | ------ |
| tyrantlucifer | 26  | male   |

#### Orc File Type[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#orc-file-type) <a href="#orc-file-type" id="orc-file-type"></a>

If you assign file type to `parquet` `orc`, schema option not required, connector can find the schema of upstream data automatically.

| Orc Data type                        | Nexus Data type                                            |
| ------------------------------------ | ---------------------------------------------------------- |
| BOOLEAN                              | BOOLEAN                                                    |
| INT                                  | INT                                                        |
| BYTE                                 | BYTE                                                       |
| SHORT                                | SHORT                                                      |
| LONG                                 | LONG                                                       |
| FLOAT                                | FLOAT                                                      |
| DOUBLE                               | DOUBLE                                                     |
| BINARY                               | BINARY                                                     |
| <p>STRING<br>VARCHAR<br>CHAR<br></p> | STRING                                                     |
| DATE                                 | LOCAL\_DATE\_TYPE                                          |
| TIMESTAMP                            | LOCAL\_DATE\_TIME\_TYPE                                    |
| DECIMAL                              | DECIMAL                                                    |
| LIST(STRING)                         | STRING\_ARRAY\_TYPE                                        |
| LIST(BOOLEAN)                        | BOOLEAN\_ARRAY\_TYPE                                       |
| LIST(TINYINT)                        | BYTE\_ARRAY\_TYPE                                          |
| LIST(SMALLINT)                       | SHORT\_ARRAY\_TYPE                                         |
| LIST(INT)                            | INT\_ARRAY\_TYPE                                           |
| LIST(BIGINT)                         | LONG\_ARRAY\_TYPE                                          |
| LIST(FLOAT)                          | FLOAT\_ARRAY\_TYPE                                         |
| LIST(DOUBLE)                         | DOUBLE\_ARRAY\_TYPE                                        |
| Map\<K,V>                            | MapType, This type of K and V will transform to Nexus type |
| STRUCT                               | NexusRowType                                               |

#### Parquet File Type[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#parquet-file-type) <a href="#parquet-file-type" id="parquet-file-type"></a>

If you assign file type to `parquet` `orc`, schema option not required, connector can find the schema of upstream data automatically.

| Orc Data type           | Nexus Data type                                            |
| ----------------------- | ---------------------------------------------------------- |
| INT\_8                  | BYTE                                                       |
| INT\_16                 | SHORT                                                      |
| DATE                    | DATE                                                       |
| TIMESTAMP\_MILLIS       | TIMESTAMP                                                  |
| INT64                   | LONG                                                       |
| INT96                   | TIMESTAMP                                                  |
| BINARY                  | BYTES                                                      |
| FLOAT                   | FLOAT                                                      |
| DOUBLE                  | DOUBLE                                                     |
| BOOLEAN                 | BOOLEAN                                                    |
| FIXED\_LEN\_BYTE\_ARRAY | <p>TIMESTAMP<br>DECIMAL</p>                                |
| DECIMAL                 | DECIMAL                                                    |
| LIST(STRING)            | STRING\_ARRAY\_TYPE                                        |
| LIST(BOOLEAN)           | BOOLEAN\_ARRAY\_TYPE                                       |
| LIST(TINYINT)           | BYTE\_ARRAY\_TYPE                                          |
| LIST(SMALLINT)          | SHORT\_ARRAY\_TYPE                                         |
| LIST(INT)               | INT\_ARRAY\_TYPE                                           |
| LIST(BIGINT)            | LONG\_ARRAY\_TYPE                                          |
| LIST(FLOAT)             | FLOAT\_ARRAY\_TYPE                                         |
| LIST(DOUBLE)            | DOUBLE\_ARRAY\_TYPE                                        |
| Map\<K,V>               | MapType, This type of K and V will transform to Nexus type |
| STRUCT                  | NexusRowType                                               |

### Options[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#options) <a href="#options" id="options"></a>

| name                            | type    | required | default value                                         | Description                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------- | ------- | -------- | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| path                            | string  | yes      | -                                                     | The s3 path that needs to be read can have sub paths, but the sub paths need to meet certain format requirements. Specific requirements can be referred to "parse\_partition\_from\_path" option                                                                                                                                                                                                           |
| file\_format\_type              | string  | yes      | -                                                     | File type, supported as the following file types: `text` `csv` `parquet` `orc` `json` `excel` `xml` `binary`                                                                                                                                                                                                                                                                                               |
| bucket                          | string  | yes      | -                                                     | The bucket address of s3 file system, for example: `s3n://nexus-test`, if you use `s3a` protocol, this parameter should be `s3a://nexus-test`.                                                                                                                                                                                                                                                             |
| fs.s3a.endpoint                 | string  | yes      | -                                                     | fs s3a endpoint                                                                                                                                                                                                                                                                                                                                                                                            |
| fs.s3a.aws.credentials.provider | string  | yes      | com.amazonaws.auth.InstanceProfileCredentialsProvider | The way to authenticate s3a. We only support `org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider` and `com.amazonaws.auth.InstanceProfileCredentialsProvider` now. More information about the credential provider you can see [Hadoop AWS Document](https://hadoop.apache.org/docs/stable/hadoop-aws/tools/hadoop-aws/index.html#Simple_name.2Fsecret_credentials_with_SimpleAWSCredentialsProvider.2A) |
| read\_columns                   | list    | no       | -                                                     | The read column list of the data source, user can use it to implement field projection. The file type supported column projection as the following shown: `text` `csv` `parquet` `orc` `json` `excel` `xml` . If the user wants to use this feature when reading `text` `json` `csv` files, the "schema" option must be configured.                                                                        |
| access\_key                     | string  | no       | -                                                     | Only used when `fs.s3a.aws.credentials.provider = org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider`                                                                                                                                                                                                                                                                                                   |
| access\_secret                  | string  | no       | -                                                     | Only used when `fs.s3a.aws.credentials.provider = org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider`                                                                                                                                                                                                                                                                                                   |
| hadoop\_s3\_properties          | map     | no       | -                                                     | If you need to add other option, you could add it here and refer to this [link](https://hadoop.apache.org/docs/stable/hadoop-aws/tools/hadoop-aws/index.html)                                                                                                                                                                                                                                              |
| delimiter/field\_delimiter      | string  | no       | \001                                                  | Field delimiter, used to tell connector how to slice and dice fields when reading text files. Default `\001`, the same as hive's default delimiter.                                                                                                                                                                                                                                                        |
| parse\_partition\_from\_path    | boolean | no       | true                                                  | Control whether parse the partition keys and values from file path. For example if you read a file from path `s3n://hadoop-cluster/tmp/nexus/parquet/name=tyrantlucifer/age=26`. Every record data from file will be added these two fields: name="tyrantlucifer", age=16                                                                                                                                  |
| date\_format                    | string  | no       | yyyy-MM-dd                                            | Date type format, used to tell connector how to convert string to date, supported as the following formats:`yyyy-MM-dd` `yyyy.MM.dd` `yyyy/MM/dd`. default `yyyy-MM-dd`                                                                                                                                                                                                                                    |
| datetime\_format                | string  | no       | yyyy-MM-dd HH:mm:ss                                   | Datetime type format, used to tell connector how to convert string to datetime, supported as the following formats:`yyyy-MM-dd HH:mm:ss` `yyyy.MM.dd HH:mm:ss` `yyyy/MM/dd HH:mm:ss` `yyyyMMddHHmmss`                                                                                                                                                                                                      |
| time\_format                    | string  | no       | HH:mm:ss                                              | Time type format, used to tell connector how to convert string to time, supported as the following formats:`HH:mm:ss` `HH:mm:ss.SSS`                                                                                                                                                                                                                                                                       |
| skip\_header\_row\_number       | long    | no       | 0                                                     | Skip the first few lines, but only for the txt and csv. For example, set like following:`skip_header_row_number = 2`. Then Nexus will skip the first 2 lines from source files                                                                                                                                                                                                                             |
| schema                          | config  | no       | -                                                     | The schema of upstream data.                                                                                                                                                                                                                                                                                                                                                                               |
| sheet\_name                     | string  | no       | -                                                     | Reader the sheet of the workbook,Only used when file\_format is excel.                                                                                                                                                                                                                                                                                                                                     |
| xml\_row\_tag                   | string  | no       | -                                                     | Specifies the tag name of the data rows within the XML file, only valid for XML files.                                                                                                                                                                                                                                                                                                                     |
| xml\_use\_attr\_format          | boolean | no       | -                                                     | Specifies whether to process data using the tag attribute format, only valid for XML files.                                                                                                                                                                                                                                                                                                                |
| compress\_codec                 | string  | no       | none                                                  |                                                                                                                                                                                                                                                                                                                                                                                                            |
| encoding                        | string  | no       | UTF-8                                                 |                                                                                                                                                                                                                                                                                                                                                                                                            |
| common-options                  |         | no       | -                                                     | Source plugin common parameters, please refer to [Source Common Options](/data-integration-with-nexus/nexus-elements/connectors/source/source-common-options.md) for details.                                                                                                                                                                                                                              |

#### delimiter/field\_delimiter \[string][​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#delimiterfield_delimiter-string) <a href="#delimiterfield_delimiter-string" id="delimiterfield_delimiter-string"></a>

**delimiter** parameter will deprecate after version 2.3.5, please use **field\_delimiter** instead.

#### compress\_codec \[string][​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#compress_codec-string) <a href="#compress_codec-string" id="compress_codec-string"></a>

The compress codec of files and the details that supported as the following shown:

* txt: `lzo` `none`
* json: `lzo` `none`
* csv: `lzo` `none`
* orc/parquet:\
  automatically recognizes the compression type, no additional settings required.

#### encoding \[string][​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#encoding-string) <a href="#encoding-string" id="encoding-string"></a>

Only used when file\_format\_type is json,text,csv,xml. The encoding of the file to read. This param will be parsed by `Charset.forName(encoding)`.

### Example[​](https://seatunnel.apache.org/docs/2.3.7/connector-v2/source/S3File#example) <a href="#example" id="example"></a>

1. In this example, We read data from s3 path `s3a://nexus-test/nexus/text` and the file type is orc in this path. We use `org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider` to authentication so `access_key` and `secret_key` is required. All columns in the file will be read and send to sink.

```
# Defining the runtime environment
env {
  parallelism = 1
  job.mode = "BATCH"
}

source {
  S3File {
    path = "/nexus/text"
    fs.s3a.endpoint="s3.cn-north-1.amazonaws.com.cn"
    fs.s3a.aws.credentials.provider = "org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider"
    access_key = "xxxxxxxxxxxxxxxxx"
    secret_key = "xxxxxxxxxxxxxxxxx"
    bucket = "s3a://nexus-test"
    file_format_type = "orc"
  }
}

transform {
  # If you would like to get more information about how to configure Nexus and see full list of transform plugins,
    # please go to transform page
}

sink {
  Console {}
}
```

2. Use `InstanceProfileCredentialsProvider` to authentication The file type in S3 is json, so need config schema option.

```

  S3File {
    path = "/nexus/json"
    bucket = "s3a://nexus-test"
    fs.s3a.endpoint="s3.cn-north-1.amazonaws.com.cn"
    fs.s3a.aws.credentials.provider="com.amazonaws.auth.InstanceProfileCredentialsProvider"
    file_format_type = "json"
    schema {
      fields {
        id = int 
        name = string
      }
    }
  }

```

3. Use `InstanceProfileCredentialsProvider` to authentication The file type in S3 is json and has five fields (`id`, `name`, `age`, `sex`, `type`), so need config schema option. In this job, we only need send `id` and `name` column to mysql.

```
# Defining the runtime environment
env {
  parallelism = 1
  job.mode = "BATCH"
}

source {
  S3File {
    path = "/nexus/json"
    bucket = "s3a://nexus-test"
    fs.s3a.endpoint="s3.cn-north-1.amazonaws.com.cn"
    fs.s3a.aws.credentials.provider="com.amazonaws.auth.InstanceProfileCredentialsProvider"
    file_format_type = "json"
    read_columns = ["id", "name"]
    schema {
      fields {
        id = int 
        name = string
        age = int
        sex = int
        type = string
      }
    }
  }
}

transform {
  # If you would like to get more information about how to configure Nexus and see full list of transform plugins,
    # please go to transoform page
}

sink {
  Console {}
}
```
