> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sync.cdata.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Microsoft OneLake

export const CommonAuthSchemeAzuremsi = ({siteName = "CData Sync"}) => {
  return <>
      <p>
        Azure 仮想マシン上で{siteName} が実行されている場合にAzure マネージドサービスID（MSI）を利用するには、<strong>Auth Scheme</strong> で<strong>Azure MSI</strong> を選択します。追加のプロパティは必要ありません。
      </p>
    </>;
};

export const CommonDatasourceAddConnector = ({datasource = "the data source", title = "the connector", destination = false, siteNameShort = "Sync"}) => {
  return <>
      <p>
        {siteNameShort} で{datasource} のデータを使用できるようにするには、まず以下の手順でコネクタを追加する必要があります：
      </p>
      <ol>
        <li>{siteNameShort} のダッシュボードから<strong>接続</strong>ページを開きます。</li>
        <li><strong>接続を追加</strong>をクリックして<strong>コネクタを選択</strong>ページを開きます。</li>
        <li>
          <strong>{destination ? "同期先" : "データソース"}</strong>タブをクリックして
          <strong>{title}</strong> 行に移動します。
        </li>
        <li>
          行末にある<strong>接続を設定</strong>アイコンをクリックして、<strong>新しい接続</strong>ページを開きます。この操作により、{' '}
          <strong>接続を追加</strong>ダイアログボックスが開きます。
          <br />
          <strong>Note：</strong> <strong>接続を設定</strong>アイコンが表示されない場合は、<strong>コネクタをダウンロード</strong>アイコンをクリックして {title} コネクタをインストールしてください。
        </li>
        <li><strong>接続を追加</strong>ダイアログボックスに接続名を入力します。</li>
        <li><strong>追加</strong>をクリックして、コネクタの<strong>設定</strong>タブを開きます。</li>
      </ol>
      <p>
        新しいコネクタのインストールについての詳細は、{' '}
        <a href="../connections">接続</a>を参照してください。
      </p>
    </>;
};

export const CommonDatasourceIntroSource = ({datasource = "the data source", siteName = "CData Sync"}) => {
  return <p>
      {siteName} アプリケーションから{datasource} コネクタを使用して、{datasource} からデータを取得してサポートされている任意の同期先に移動できます。これを行うには、コネクタを追加し、コネクタへの認証を行い、接続を完了する必要があります。
    </p>;
};

export const driverVersion = "M";

export const siteNameShort = "Sync";

export const siteName = "CData Sync";

export const datasource = "Microsoft OneLake";
export const pageTitle = "Microsoft OneLake";
export const oauthschemes = "AzureAD";

<CommonDatasourceIntroSource datasource={datasource} siteName={siteName} />

## Supported File Formats

When {siteNameShort} writes data to {datasource}, you can choose the file format for the exported data. The following file formats are supported for the {datasource} destination:

* **CSV：** Plain text comma-separated values.
* **Avro：** A row-based binary format that supports schema evolution.
* (Default) **Parquet：** A columnar storage format that is optimized for analytics.

## Microsoft OneLake コネクタを追加

<CommonDatasourceAddConnector datasource={datasource} title={pageTitle} destination={false} siteNameShort={siteNameShort} />

## Microsoft OneLake への認証

After you add the connector, you need to set the required properties.

* **Connection Name:** Enter a connection name of your choice.
* **File Format:** Select the file format that you want to use: **CSV** (default), **Avro**, and **Parquet**.
* **Azure Storage Account:** Enter the name of your Azure storage account.
* **URI:** Enter the path of the file system and folder that contains your files (for example, `onelake://Workspace/Test.LakeHouse/Files/CustomFolder`).

{siteName} supports authenticating to {pageTitle} in several ways. Select your authentication method below to proceed to the relevant section that contains the authentication details.

* [**Azure Active Directory**](#azure-active-directory) (default)
* [**Azure Managed Service Identity**](#azure-managed-service-identity)
* [**Azure Service Principal**](#azure-service-principal)
* [**Azure Service Principal Certificate**](#azure-service-principal-certificate)

### Azure Active Directory

To connect with an Azure Active Directory (AD) user account, specify the following properties:

* **Auth Scheme:** Select **AzureAD**.
* **Use Lake Formation:** Select **True** if you want the AWS Lake Formation service to retrieve temporary credentials. These temporary credentials enforce access policies against the user based on the configured IAM role. You can use this service when you authenticate through AzureAD, Okta, ADFS, and PingFederate, while providing a Security Assertion Markup Language (SAML) assertion. The default setting for **Use Lake Formation** is **False**.

### Azure Managed Service Identity

<CommonAuthSchemeAzuremsi siteName={siteName} siteNameShort={siteNameShort} datasource={datasource} />

## 接続を完了する

To complete your connection:

1. Specify the following properties:

   **For the CSV file format:**

   * **FMT:** Enter the format that you want to use to parse all text files. The default format is **CsvDelimited**.
   * **Aggregate Files:** Specify whether you want to aggregate all the files that are located in the URI directory and that have the same schema into a single table named AggregatedFiles. The default option is **False**.
   * **Include Column Headers:** Specify whether you want to obtain column headers from the first lines of the specified files. The default option is **True.**

   **For the Avro and Parquet file formats:**

   * **Data Model:** Select the data model that you want to use to parse documents for your format and to generate the database metadata. The default data model is **Document**.
   * **Aggregate Files:** Specify whether you want to aggregate all the files that are located in the URI directory and that have the same schema into a single table named AggregatedFiles. The default option is **False**.
2. Define advanced connection settings on the **Advanced** tab. (In most cases, though, you should not need these settings.)
3. {oauthschemes ? `If you authenticate with ${oauthschemes}, click` : "Click"} **Connect to {datasource}** to connect to your {pageTitle} account.
4. Click **Create & Test** to create your connection.

## **Load Folder** Job Behavior for Parquet and Avro Files

{siteNameShort} optimizes **Load Folder** jobs that use Microsoft OneLake sources with Parquet or Avro files to reduce redundant downloads and improve performance.

During a **Load Folder** job run, {siteNameShort} downloads each file once and reuses it for both schema detection and data processing. This behavior reduces overall data transfer and improves performance for large datasets or folders with many files.

Temporary files are stored in a job-specific directory:

```
{ApplicationDatabase}/connections/<SourceConnection>/temp/<JobId>/
```

In this path:

* `{ApplicationDatabase}` specifies the root directory of the {siteNameShort} application database.
* `<SourceConnection>` specifies the name or identifier of the source connection.
* `<JobId>` specifies the system-generated job identifier.

Temporary files are removed automatically after the job completes, regardless of success or failure.

<Note>Ensure that sufficient disk space is available for the downloaded files during job execution.</Note>
