> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sync.cdata.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Azure Data Lake Storage

export const PrereqsDeltaFileFormats = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <h2>Prerequisites</h2>
      <p>
        Before you configure the {datasource} destination with the <strong>Delta Parquet</strong> file format in the Microsoft Windows operating system (OS), make sure that your environment meets the requirements explained below. These prerequisites ensure that {siteNameShort} can interact correctly with Delta Lake by locating the required Hadoop binaries under Windows.
      </p>
      <p>Configure your Windows OS, as follows:</p>
      <ul>
        <li>Download Hadoop binaries (recommended version: 2.8.1 or later).</li>
        <li>Ensure that <code>HADOOP_HOME</code> environment variable to the Hadoop installation directory.</li>
        <li>Ensure that <code>%HADOOP_HOME%\bin</code> is included in your <code>PATH</code> system variable (specifically, <code>%HADOOP_HOME%\bin\winutils.exe</code> must be accessible).</li>
      </ul>
      <p>
        This configuration is necessary because Delta Lake (on Spark) uses Hadoop's file system APIs to access local storage, the Hadoop Distributed File System (HDFS), and cloud object stores like {datasource}. Under Windows, Spark must be able to locate the Hadoop binaries (including winutils.exe and other native libraries) to function correctly. Without this configuration, operations such as writing Delta tables, managing checkpoints, or accessing cloud storage can fail with permission or file-system errors.
      </p>
    </>;
};

export const CommonAuthSchemeAzurestoragesas = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>Azure の共有アクセス署名で接続するには、次のプロパティを設定します：</p>
      <ul>
        <li><strong>Auth Scheme：</strong><strong>AzureStorageSAS</strong> を選択。</li>
        <li><strong>Azure Shared Access Signature：</strong>ストレージアカウントに関連付けられている共有アクセス署名を入力。</li>
      </ul>
      <p>Azure 共有アクセス署名を作成するには：</p>
      <ol>
        <li><a href="https://portal.azure.com" target="_blank" rel="noopener noreferrer">Azure ポータル</a>にルートアカウントの資格情報を使用してサインインします。</li>
        <li><strong>ストレージ アカウント</strong>をクリックして、使用するストレージアカウントを選択します。</li>
        <li><strong>設定</strong>の<strong>Shared Access Signature</strong> をクリックします。</li>
        <li>アクセスポリシーとトークンの有効期限を設定します。</li>
        <li><strong>SAS の生成</strong>をクリックし、生成されたトークンをコピーします。</li>
      </ol>
    </>;
};

export const CommonAuthSchemeAccesskey = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>Azure のアクセスキーで接続するには、次のプロパティを設定します：</p>
      <ul>
        <li><strong>Auth Scheme：</strong><strong>Access Key</strong> を選択。</li>
        <li><strong>Azure Access Key：</strong>ストレージアカウントに関連付けられているアクセスキーを入力。</li>
      </ul>
      <p>アクセスキーを取得するには：</p>
      <ol>
        <li><a href="https://portal.azure.com/" target="_blank" rel="noopener noreferrer">Azure ポータル</a>にルートアカウントの資格情報を使用してサインインします。</li>
        <li><strong>ストレージ アカウント</strong>をクリックして、使用するストレージアカウントを選択します。</li>
        <li><strong>設定</strong>の<strong>アクセス キー</strong>をクリックします。ページ上にストレージ アカウント名とキーが表示されます。</li>
      </ol>
    </>;
};

export const CommonAuthSchemeAzureserviceprincipalcertFileFormat = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>To connect with an Azure service principal and client certificate, set the following properties:</p>
      <ul>
        <li><strong>Auth Scheme:</strong> Select <strong>AzureServicePrincipalCert</strong>.</li>
        <li><strong>Azure Tenant:</strong> Enter the Microsoft Online tenant to which you want to connect.</li>
        <li><strong>OAuth JWT Cert:</strong> Enter your Java web tokens (JWT) certificate store.</li>
        <li><strong>OAuth JWT Cert Type:</strong> Enter the type of key store that contains your JWT Certificate. The default type is <strong>PEMKEY_BLOB</strong>.</li>
        <li><strong>OAuth Client Id:</strong> Enter the client Id that you were assigned when you registered your application with an OAuth authorization server.</li>
        <li>(Optional) <strong>Scope:</strong> Specify the scope of your access to the application.</li>
        <li>(Optional) <strong>OAuth Authorization URL:</strong> Enter the OAuth authorization URL for the OAuth service.</li>
        <li>(Optional) <strong>OAuth Access Token URL:</strong> Enter the URL from which to retrieve the access token.</li>
        <li>(Optional) <strong>OAuth Refresh Token URL:</strong> Enter the URL from which to refresh the OAuth token.</li>
        <li>(Optional) <strong>OAuth JWT Cert Password:</strong> Enter the password for your OAuth JWT certificate.</li>
        <li>(Optional) <strong>OAuth JWT Cert Subject:</strong> Enter the subject of your OAuth JWT certificate.</li>
      </ul>
      <p>To obtain the OAuth certificate for your application:</p>
      <ol>
        <li>Log in to the <a href="https://portal.azure.com" target="_blank" rel="noopener noreferrer">Azure portal</a>.</li>
        <li>In the left navigation pane, select <strong>All services</strong>. Then, search for and select <strong>App registrations</strong>.</li>
        <li>Click <strong>New registrations</strong>.</li>
        <li>Enter an application name and select <strong>Any Azure AD Directory - Multi Tenant</strong>.</li>
        <li>After you create the application, copy the application (client) Id value that is displayed in the <strong>Overview</strong> section. Use this value as the OAuth client Id.</li>
        <li>Navigate to the <strong>Certificates & Secrets</strong> section and select <strong>Upload certificate</strong>. Then, select the certificate to upload from your local machine.</li>
        <li>Specify the duration and save the client secret. After you save it, the key value is displayed.</li>
        <li>Copy this value because it is displayed only once. You will use this value as the OAuth client secret.</li>
        <li>On the <strong>Authentication</strong> tab, make sure to select <strong>Access tokens (used for implicit flows)</strong>.</li>
      </ol>
    </>;
};

export const CommonAuthSchemeAzureserviceprincipalFileFormat = ({siteName = "CData Sync", siteNameShort = "Sync", datasource = "the data source"}) => {
  return <>
      <p>To connect with an Azure service principal and client secret, set the following properties:</p>
      <ul>
        <li><strong>Auth Scheme:</strong> Select <strong>AzureServicePrincipal</strong>.</li>
        <li><strong>Azure Tenant:</strong> Enter the Microsoft Online tenant to which you want to connect.</li>
        <li><strong>OAuth Client Id:</strong> Enter the client Id that you were assigned when you registered your application with an OAuth authorization server.</li>
        <li><strong>OAuth Client Secret:</strong> Enter the client secret that you were assigned when you registered your application with an OAuth authorization server.</li>
        <li>(Optional) <strong>Scope:</strong> Specify the scope of your access to the application.</li>
        <li>(Optional) <strong>OAuth Authorization URL:</strong> Enter the OAuth authorization URL for the OAuth service.</li>
        <li>(Optional) <strong>OAuth Access Token URL:</strong> Enter the URL from which to retrieve the access token.</li>
        <li>(Optional) <strong>OAuth Refresh Token URL:</strong> Enter the URL from which to refresh the OAuth token.</li>
      </ul>
      <p>To obtain the OAuth client Id and client secret for your application:</p>
      <ol>
        <li>Log in to the <a href="https://portal.azure.com" target="_blank" rel="noopener noreferrer">Azure portal</a>.</li>
        <li>In the left navigation pane, select <strong>All services</strong>. Then, search for and select <strong>App registrations</strong>.</li>
        <li>Click <strong>New registrations</strong>.</li>
        <li>Enter an application name and select <strong>Any Azure AD Directory - Multi Tenant</strong>.</li>
        <li>After you create the application, copy the application (client) Id value that is displayed in the <strong>Overview</strong> section. Use this value as the OAuth client Id.</li>
        <li>Navigate to the <strong>Certificates & Secrets</strong> section and select <strong>New Client Secret</strong> for the application.</li>
        <li>Specify the duration and save the client secret. After you save it, the key value is displayed.</li>
        <li>Copy this value because it is displayed only once. You will use this value as the OAuth client secret.</li>
        <li>On the <strong>Authentication</strong> tab, make sure to select <strong>Access tokens (used for implicit flows)</strong>.</li>
      </ol>
    </>;
};

export const CommonAuthSchemeAzuremsi = ({siteName = "CData Sync"}) => {
  return <>
      <p>
        Azure 仮想マシン上で{siteName} が実行されている場合にAzure マネージドサービスID（MSI）を利用するには、<strong>Auth Scheme</strong> で<strong>Azure MSI</strong> を選択します。追加のプロパティは必要ありません。
      </p>
    </>;
};

export const CommonDatasourceAddConnector = ({datasource = "the data source", title = "the connector", destination = false, siteNameShort = "Sync"}) => {
  return <>
      <p>
        {siteNameShort} で{datasource} のデータを使用できるようにするには、まず以下の手順でコネクタを追加する必要があります：
      </p>
      <ol>
        <li>{siteNameShort} のダッシュボードから<strong>接続</strong>ページを開きます。</li>
        <li><strong>接続を追加</strong>をクリックして<strong>コネクタを選択</strong>ページを開きます。</li>
        <li>
          <strong>{destination ? "同期先" : "データソース"}</strong>タブをクリックして
          <strong>{title}</strong> 行に移動します。
        </li>
        <li>
          行末にある<strong>接続を設定</strong>アイコンをクリックして、<strong>新しい接続</strong>ページを開きます。この操作により、{' '}
          <strong>接続を追加</strong>ダイアログボックスが開きます。
          <br />
          <strong>Note：</strong> <strong>接続を設定</strong>アイコンが表示されない場合は、<strong>コネクタをダウンロード</strong>アイコンをクリックして {title} コネクタをインストールしてください。
        </li>
        <li><strong>接続を追加</strong>ダイアログボックスに接続名を入力します。</li>
        <li><strong>追加</strong>をクリックして、コネクタの<strong>設定</strong>タブを開きます。</li>
      </ol>
      <p>
        新しいコネクタのインストールについての詳細は、{' '}
        <a href="../connections">接続</a>を参照してください。
      </p>
    </>;
};

export const CommonDatasourceIntroDestination = ({datasource = "the data source", siteName = "CData Sync"}) => {
  return <p>
      {siteName} アプリケーションから{datasource} コネクタを使用して、サポートされている任意のデータソースから同期先の{datasource} へデータを移動できます。これを行うには、コネクタを追加し、コネクタへの認証を行い、接続を完了する必要があります。
    </p>;
};

export const driverVersion = "M";

export const siteNameShort = "Sync";

export const siteName = "CData Sync";

<CommonDatasourceIntroDestination datasource="Azure Data Lake Storage" siteName={siteName} />

<PrereqsDeltaFileFormats datasource="Azure Data Lake Storage" siteNameShort={siteNameShort} />

## Supported File Formats

When {siteNameShort} writes data to Azure Data Lake Storage, you can choose the file format for the exported data. The following file formats are supported for the Azure Data Lake Storage destination:

* (Default) **Delta Parquet：** A format that uses a Delta Lake storage layer on top of the Parquet file format that is used by {siteNameShort} to support delta processing. *Delta processing* is a method where, after your initial job run, only new or modified files are written or read in subsequent runs, which can reduce job times and resource use.

  **Limitations:**

  * **Naming restrictions:** Table and column names cannot include special characters or reserved SQL and Delta Lake keywords. Examples of special characters include spaces, commas, semicolons, braces, parentheses, equal signs, and the newline (`\n`) and tab (`\t`) characters.
  * **Primary keys:** Primary key constraints are not supported. {siteNameShort} uses the source primary keys for incremental replication.
  * **Data types:** Unlike traditional databases, Delta Lake does not support column-size definitions (for example, `VARCHAR(100)`). It supports only a fixed set of data types and allows type widening when necessary.
  * **Schema changes:** The ALTER TABLE command supports only adding new columns. Changing the data type of an existing column (for example, from INT to VARCHAR) is not supported.
  * **Delete operations:** In standard jobs, both hard and soft deletions are supported. In CDC and enhanced CDC jobs, only soft deletions are supported.
* **Apache Iceberg：** A high-performance table format that supports atomicity, consistency, isolation, and durability (ACID) transactions and schema evolution.
* **CSV：** Plain text comma-separated values.
* **Avro：** A row-based binary format that supports schema evolution.
* **Parquet：** A columnar storage format that is optimized for analytics.

## Azure Data Lake Storage コネクタを追加

<CommonDatasourceAddConnector datasource="Azure Data Lake Storage" title="Azure Data Lake Storage" destination={true} siteNameShort={siteNameShort} />

## Azure Data Lake Storage への認証

After you add the connector, you need to set the required properties.

* **File Format:** Select the file format that you want to use: **Delta Parquet**, **Apache Iceberg** (default), **CSV**, **Avro**, or **Parquet**.
* **URI:** Enter the path of the file system and folder that contains your files (for example, `abfss://MyFileSystem/FolderName`).
* **Azure Storage Account:** Enter the name of your Azure storage account.

{siteName} supports authenticating to Azure Data Lake Storage in several ways. Select your authentication method below to proceed to the relevant section that contains the authentication details.

* [**Azure Active Directory**](#azure-active-directory) (default)
* [**Azure Managed Service Identity**](#azure-managed-service-identity)
* [**Azure Service Principal**](#azure-service-principal)
* [**Azure Service Principal Certificate**](#azure-service-principal-certificate)
* [**Access Key**](#access-key)
* [**Azure Storage SAS**](#azure-shared-access-signature)

**Note:** The full list of authentication methods above are for all file formats except **Delta Parquet**. That format uses only the **Access Key** method.

### Azure Active Directory

To connect with an Azure Active Directory (AD) user account, specify the following properties:

* **Auth Scheme:** Select **AzureAD**.
* **Use Lake Formation:** Select whether you want the AWS Lake Formation service to retrieve temporary credentials. These temporary credentials enforce access policies against the user based on the configured IAM role. You can use this service when you authenticate through AzureAD, Okta, ADFS, and PingFederate, while providing a Security Assertion Markup Language (SAML) assertion. By default, the **Enable** checkbox is not selected.

### Azure Managed Service Identity

<CommonAuthSchemeAzuremsi siteName={siteName} />

### Azure Service Principal

<CommonAuthSchemeAzureserviceprincipalFileFormat siteName={siteName} />

### Azure Service Principal Certificate

<CommonAuthSchemeAzureserviceprincipalcertFileFormat siteName={siteName} />

### Azure Access Key

<CommonAuthSchemeAccesskey siteName={siteName} />

### Azure Shared Access Signature

<CommonAuthSchemeAzurestoragesas siteName={siteName} />

## 接続を完了する

To complete your connection:

1. Specify the following properties:

   **For the Delta Parquet and CSV file formats:**

   * **FMT:** Enter the format that you want to use to parse all text files. The default format is **CsvDelimited**.
   * **Aggregate Files:** Select whether you want to aggregate all the files that are located in the URI directory and that have the same schema into a single table named AggregatedFiles. By default, the **Enable** checkbox is not selected.
   * **Include Column Headers:** Select whether you want to obtain column headers from the first lines of the specified files. By default, the **Enable** checkbox is already selected.

   **For the Avro and Parquet file formats:**

   * **Data Model:** Select the data model that you want to use to parse documents for your format and to generate the database metadata. The default data model is **Document**.
   * **Aggregate Files:** Select whether you want to aggregate all the files that are located in the URI directory and that have the same schema into a single table named AggregatedFiles. By default, the **Enable** checkbox is not selected.
2. Define advanced connection settings on the **Advanced** tab. (In most cases, though, you should not need these settings.)
3. If you authenticate with AzureAD, click **Connect to Azure Data Lake Storage** to connect to your Azure Data Lake Storage account.
4. Click **Create & Test** to create your connection.
