> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sync.cdata.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Amazon S3

export const PrereqsDeltaFileFormats = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <h2>Prerequisites</h2>
      <p>
        Before you configure the {datasource} destination with the <strong>Delta Parquet</strong> file format in the Microsoft Windows operating system (OS), make sure that your environment meets the requirements explained below. These prerequisites ensure that {siteNameShort} can interact correctly with Delta Lake by locating the required Hadoop binaries under Windows.
      </p>
      <p>Configure your Windows OS, as follows:</p>
      <ul>
        <li>Download Hadoop binaries (recommended version: 2.8.1 or later).</li>
        <li>Ensure that <code>HADOOP_HOME</code> environment variable to the Hadoop installation directory.</li>
        <li>Ensure that <code>%HADOOP_HOME%\bin</code> is included in your <code>PATH</code> system variable (specifically, <code>%HADOOP_HOME%\bin\winutils.exe</code> must be accessible).</li>
      </ul>
      <p>
        This configuration is necessary because Delta Lake (on Spark) uses Hadoop's file system APIs to access local storage, the Hadoop Distributed File System (HDFS), and cloud object stores like {datasource}. Under Windows, Spark must be able to locate the Hadoop binaries (including winutils.exe and other native libraries) to function correctly. Without this configuration, operations such as writing Delta tables, managing checkpoints, or accessing cloud storage can fail with permission or file-system errors.
      </p>
    </>;
};

export const CommonAuthSchemeAwscredentialsfile = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>To connect with a credentials file, specify the following properties:</p>
      <ul>
        <li><strong>Auth Scheme:</strong> Select <strong>AwsCredentialsFile</strong>.</li>
        <li><strong>AWS Credentials File:</strong> Enter the location of your Amazon Web Services (AWS) credentials file.</li>
        <li>(Optional) <strong>AWS Credentials File Profile:</strong> Enter the name of the AWS profile that you want to use from the credentials file that you specify. If you do not enter a profile name, {siteNameShort} uses the profile named <strong>default</strong>.</li>
      </ul>
    </>;
};

export const CommonAuthSchemePingfederateSsoFileFormat = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <ul>
        <li><strong>Auth Scheme:</strong> Select <strong>PingFederate</strong>.</li>
        <li><strong>User:</strong> Enter the username that you use to authenticate to your PingFederate account.</li>
        <li><strong>Password:</strong> Enter the password that you use to authenticate to your PingFederate account.</li>
        <li><strong>SSO Login URL</strong> Enter the login URL that is used by your SSO provider.</li>
        <li><strong>SSO Exchange UrI:</strong> Enter the Partner Service Identifier URI that is configured in your PingFederate server instance. The URI is available under <strong>SP Connections &gt; SP Connection &gt; WS-Trust &gt; Protocol Settings</strong>.</li>
        <li><strong>Use Lake Formation:</strong> Select <strong>True</strong> if you want the AWS Lake Formation service to retrieve temporary credentials. These temporary credentials enforce access policies against the user based on the configured IAM role. You can use this service when you authenticate through AzureAD, Okta, ADFS, and PingFederate, while providing a Security Assertion Markup Language (SAML) assertion. The default setting for <strong>Use Lake Formation</strong> is <strong>False</strong>.</li>
        <li>(Optional) <strong>AWS Principal ARN:</strong> The Amazon Resource Name (ARN) of the Security Assertion Markup Language (SAML) identity provider in your AWS account.</li>
        <li>(Optional) <strong>SSO Properties:</strong> Enter a comma-separated list of the single sign-on (SSO) properties that you want to use (for example, <em>SSOProperty1</em>=<em>Value1</em>;<em>SSOProperty2</em>=<em>Value2</em>;...).</li>
      </ul>
    </>;
};

export const CommonAuthSchemeAwsiamroles = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>To connect with your IAM user credentials, specify the following properties:</p>
      <ul>
        <li><strong>Auth Scheme:</strong> Select <strong>AwsIAMRoles</strong>.</li>
        <li><strong>User:</strong> Enter the username that you use to authenticate to your {datasource} account.</li>
        <li><strong>AWS Access Key:</strong> Enter your Amazon Web Services (AWS) account access key. You can locate this value on your AWS security credentials page.</li>
        <li><strong>AWS Secret Key:</strong> Enter your AWS account secret key. You can locate this value on your AWS security credentials page.</li>
        <li><strong>AWS Role ARN:</strong> Enter the Amazon Resource Name (ARN) for the role with which you want to authenticate.</li>
        <li>(Optional) <strong>AWS External Id:</strong> Enter the unique identifier that is required when you assume a role in another account.</li>
      </ul>
    </>;
};

export const CommonAuthSchemeAwsec2roles = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>
        When you run {siteName} on an EC2 instance, {siteName} can authenticate by using the IAM role that is assigned to the instance. Select <strong>AwsEC2Roles</strong> for <strong>Auth Scheme</strong> to use that role. No additional properties are required.
      </p>
    </>;
};

export const CommonAuthSchemeAwsrootkeys = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>To connect with your account root credentials, specify the following properties:</p>
      <ul>
        <li><strong>Auth Scheme:</strong> Select <strong>AwsRootKeys</strong>.</li>
        <li><strong>AWS Access Key:</strong> Enter your Amazon Web Services (AWS) account access key. You can locate this value on your AWS security credentials page.</li>
        <li><strong>AWS Secret Key:</strong> Enter your AWS account secret key. You can locate this value on your AWS security credentials page.</li>
      </ul>
    </>;
};

export const CommonDatasourceAddConnector = ({datasource = "the data source", title = "the connector", destination = false, siteNameShort = "Sync"}) => {
  return <>
      <p>
        To enable {siteNameShort} to use data from {datasource}, you first must add the
        connector, as follows:
      </p>
      <ol>
        <li>Open the <strong>Connections</strong> page of the {siteNameShort} dashboard.</li>
        <li>Click <strong>Add Connection</strong> to open the <strong>Select Connectors</strong> page.</li>
        <li>
          Click the <strong>{destination ? "Destinations" : "Sources"}</strong> tab and locate
          the <strong>{title}</strong> row.
        </li>
        <li>
          Click the <strong>Configure Connection</strong> icon at the end of that row to open
          the <strong>New Connection</strong> page. This action opens the{' '}
          <strong>Add Connection</strong> dialog box.
          <br />
          <strong>Note:</strong> If the <strong>Configure Connection</strong> icon is not
          available, click the <strong>Download Connector</strong> icon to install
          the {title} connector.
        </li>
        <li>Enter a name for your connection in the <strong>Add Connection</strong> dialog box.</li>
        <li>Click <strong>Add</strong> to open the <strong>Settings</strong> tab for your connector.</li>
      </ol>
      <p>
        For more information about installing new connectors, see{' '}
        <a href="../connections">Connections</a>.
      </p>
    </>;
};

export const CommonDatasourceIntroDestination = ({datasource = "the data source", siteName = "CData Sync"}) => {
  return <p>
      You can use the {datasource} connector from the {siteName} application to move data
      from any supported source to the {datasource} destination. To do so, you need to add
      the connector, authenticate to the connector, and complete your connection.
    </p>;
};

export const driverVersion = "M";

export const siteNameShort = "Sync";

export const siteName = "CData Sync";

<CommonDatasourceIntroDestination datasource="Amazon S3" siteName={siteName} />

<PrereqsDeltaFileFormats datasource="Amazon S3" siteNameShort={siteNameShort} />

## Supported File Formats

When {siteNameShort} writes data to Amazon S3, you can choose the file format for the exported data. The following file formats are supported for the Amazon S3 destination:

* (Default) **Delta Parquet:** A format that uses a Delta Lake storage layer on top of the Parquet file format that is used by {siteNameShort} to support delta processing. *Delta processing* is a method where, after your initial job run, only new or modified files are written or read in subsequent runs, which can reduce job times and resource use.

  **Limitations:**

  * **Naming restrictions:** Table and column names cannot include special characters or reserved SQL and Delta Lake keywords. Examples of special characters include spaces, commas, semicolons, braces, parentheses, equal signs, and the newline (`\n`) and tab (`\t`) characters.
  * **Primary keys:** Primary key constraints are not supported. {siteNameShort} uses the source primary keys for incremental replication.
  * **Data types:** Unlike traditional databases, Delta Lake does not support column-size definitions (for example, `VARCHAR(100)`). It supports only a fixed set of data types and allows type widening when necessary.
  * **Schema changes:** The ALTER TABLE command supports only adding new columns. Changing the data type of an existing column (for example, from INT to VARCHAR) is not supported.
  * **Delete operations:** In standard jobs, both hard and soft deletions are supported. In CDC and enhanced CDC jobs, only soft deletions are supported.
* **Apache Iceberg:** A high-performance table format that supports atomicity, consistency, isolation, and durability (ACID) transactions and schema evolution.
* **CSV:** Plain text comma-separated values.
* **Avro:** A row-based binary format that supports schema evolution.
* **Parquet:** A columnar storage format that is optimized for analytics.

## Add the Amazon S3 Connector

<CommonDatasourceAddConnector datasource="Amazon S3" title="Amazon S3" destination={true} siteNameShort={siteNameShort} />

## Authenticate to Amazon S3

After you add the connector, you need to set the required properties.

* **File Format:** Select the file format that you want to use: **Delta Parquet**, **Apache Iceberg**,**CSV** (default), **Avro**, or **Parquet**.
* **URI:** Enter the path of your bucket and folder (for example, `s3://BucketName/FolderName`).
* **AWS Region:** Select the hosting region for Amazon Web Services. The default region is **NORTHERNVIRGINIA**.

{siteName} supports authenticating to Amazon S3 in several ways, based on the file format that you select.

* [**AWS Root Keys**](#aws-root-keys) (default)
* [**AWS EC2 Roles**](#aws-ec2-roles)
* [**AWS IAM Roles**](#aws-iam-roles)
* [**Active Directory Federation Services**](#active-directory-federation-services)
* [**Okta**](#okta)
* [**PingFederate**](#pingfederate)
* [**AWS Temporary Credentials**](#aws-temporary-credentials)
* [**AWS Credentials File**](#aws-credentials-file)
* [**Azure Active Directory**](#azure-active-directory)

**Note:** The full list of authentication methods above are for all file formats except **Delta Parquet**. That format uses only the **AWS Root Keys** method.

### AWS Root Keys

<CommonAuthSchemeAwsrootkeys siteName={siteName} />

* (Optional) **MFA Serial Number:** Enter the serial number for your multifactor authentication (MFA) device, if you are using such a device.
* (Optional) **MFA Token:** Enter the temporary token that is available from your MFA device.
* **Temporary Token Duration:** Enter the duration, in seconds, that you want for your temporary credentials. The default duration is **3600**.

### AWS EC2 Roles

<CommonAuthSchemeAwsec2roles siteName={siteName} />

### AWS IAM Roles

<CommonAuthSchemeAwsiamroles siteName={siteName} datasource="Amazon S3" />

* (Optional) **MFA Serial Number:** Enter the serial number for your multifactor authentication (MFA) device, if you are using such a device.
* (Optional) **MFA Token:** Enter the temporary token that is available from your MFA device.
* **Temporary Token Duration:** Enter the duration, in seconds, that you want for your temporary credentials. The default duration is **3600**.

### Active Directory Federation Services

To connect with single sign-on (SSO) via Active Directory Federation Services (ADFS), specify the following properties:

* **Auth Scheme:** Select **ADFS**.
* **User:** Enter the username that you use to authenticate to your ADFS account.
* **Password:** Enter the password that you use to authenticate to your ADFS account.
* **SSO Login URL:** Enter the login URL that is used by your SSO provider.
* **Use Lake Formation:** Select whether you want the AWS Lake Formation service to retrieve temporary credentials. These temporary credentials enforce access policies against the user based on the configured IAM role. You can use this service when you authenticate through AzureAD, Okta, ADFS, and PingFederate, while providing a Security Assertion Markup Language (SAML) assertion. By default, the **Enable** checkbox is not selected.
* (Optional) **SSO Properties:** Enter a semicolon-separated list of the single sign-on (SSO) properties that you want to use (for example, SSOProperty1=Value1;SSOProperty2=Value2; ...).

### Okta

To connect with single sign-on (SSO) via Okta, specify the following properties:

* **Auth Scheme:** Select **Okta**.
* **User:** Enter the username that you use to authenticate to your Okta account.
* **Password:** Enter the password that you use to authenticate to your Okta account.
* **SSO Login URL:** Enter the login URL that is used by your SSO provider.
* **Use Lake Formation:** Select whether you want the AWS Lake Formation service to retrieve temporary credentials. These temporary credentials enforce access policies against the user based on the configured IAM role. You can use this service when you authenticate through AzureAD, Okta, ADFS, and PingFederate, while providing a Security Assertion Markup Language (SAML) assertion. By default, the **Enable** checkbox is not selected.
* (Optional) **SSO Properties:** Enter a semicolon-separated list of the single sign-on (SSO) properties that you want to use (for example, *SSOProperty1*=*Value1*;*SSOProperty2*=*Value2*;...).

### PingFederate

<CommonAuthSchemePingfederateSsoFileFormat siteName={siteName} />

### AWS Temporary Credentials

To connect with AWS temporary credentials, specify the following properties:

* **Auth Scheme:** Select **AwsTempCredentials**.
* **AWS Access Key:** Enter the access key that is associated with your Amazon Web Services (AWS) account. This value is accessible from your AWS security credentials page.
* **AWS Secret Key:** Enter the secret key that is associated with your AWS account. This value is accessible from your AWS security credentials page.
* **AWS Session Token:** Enter your AWS session token. This token is provided with your temporary credentials. For more information, see [AWS Identity and Access Management: User Guide](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_credentials_temp_use-resources.html).

### AWS Credentials File

<CommonAuthSchemeAwscredentialsfile siteName={siteName} siteNameShort={siteNameShort} />

### Azure Active Directory

To connect with an Azure Active Directory (AD) user account, specify the following properties:

* **Auth Scheme:** Select **AzureAD**.
* **Use Lake Formation:** Select whether you want the AWS Lake Formation service to retrieve temporary credentials. These temporary credentials enforce access policies against the user based on the configured IAM role. You can use this service when you authenticate through AzureAD, Okta, ADFS, and PingFederate, while providing a Security Assertion Markup Language (SAML) assertion. By default, the **Enable** checkbox is not selected.
* **OAuth Client Id:** Enter the client Id that you were assigned when you registered your application with an OAuth authorization server.
* **OAuth Client Secret:** Enter the client secret that you were assigned when you registered your application with an OAuth authorization server.

## Complete Your Connection

To complete your connection:

1. Specify the following properties:

   **For all formats:**

   (Optional) **Storage Base URL:** Enter the URL of your cloud-storage service provider.

   **For the Delta Parquet and CSV file formats only:**

   * **FMT:** Enter the format that you want to use to parse all text files. The default format is **CsvDelimited**.
   * **Aggregate Files:** Specify whether you want to aggregate all the files that are located in the URI directory and that have the same schema into a single table named AggregatedFiles. The default option is **False**.
   * **Include Column Headers:** Select whether you want to obtain column headers from the first lines of the specified files. By default, the **Enable** checkbox is already selected.

   **For all file formats except Delta Parquet:**

   * **Data Model:** Select the data model that you want to use to parse documents for your format and to generate the database metadata. The default data model is **Document**.
   * **Aggregate Files:** Select whether you want to aggregate all the files that are located in the URI directory and that have the same schema into a single table named AggregatedFiles. By default, the **Enable** checkbox is not selected.
2. Define advanced connection settings on the **Advanced** tab. (In most cases, though, you should not need these settings.)
3. Click **Create & Test** to create your connection.
