> For the complete documentation index, see [llms.txt](https://dataflint.gitbook.io/dataflint-for-spark/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://dataflint.gitbook.io/dataflint-for-spark/saas-platform/emr-saas-installation.md).

# Amazon EMR

## Summary

DataFlint reads your EMR job history and Spark event logs through a cross-account IAM role that you create and control. The role is **read-only**: DataFlint cannot start, stop or modify clusters, applications or jobs.

EMR on EC2, EMR Serverless and EMR on EKS are separate AWS services with separate APIs, so each has its own small policy below. Pick the tab that matches your deployment. If you run more than one, create one role per service, or combine the statements into one role.

You will:

1. Create an IAM policy with the read-only actions for your EMR service.
2. Create an IAM role that trusts the DataFlint principal with your external ID.
3. Add the role in the DataFlint console.

The entire process should take a few minutes.

{% hint style="info" %}
The DataFlint console can generate the Terraform or CloudFormation for this role with your account's values filled in (Admin Panel → Environment Management → Add New → EMR → "Need to create this role?"). The listings on this page are the reference your security team can review; if the generated file and this page differ, this page is authoritative.
{% endhint %}

### Identity and external ID

* **DataFlint principal:** `arn:aws:iam::975050001706:role/eks-dataflint-service-role`
* **Your external ID:** generated for your DataFlint account and shown in the console when you add an environment. It is required in the trust policy so that only your DataFlint account can use the role.

{% hint style="warning" %}
BYOC customers: the principal is the `eks-dataflint-service-role` in your own DataFlint account, not the one above. The console shows the correct value.
{% endhint %}

### Trust policy (all EMR services)

Replace `YOUR_EXTERNAL_ID` with the value shown in the console.

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": {
        "AWS": "arn:aws:iam::975050001706:role/eks-dataflint-service-role"
      },
      "Action": "sts:AssumeRole",
      "Condition": {
        "StringEquals": {
          "sts:ExternalId": "YOUR_EXTERNAL_ID"
        }
      }
    }
  ]
}
```

## Permissions per EMR service

{% tabs %}
{% tab title="EMR on EC2" %}

### Why each permission is needed

| Actions                                                                                                                                                                          | What DataFlint does with it                                                                                                                                                                                                                                                                  | What it cannot do                                                            |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| `elasticmapreduce:ListClusters`, `DescribeCluster`, `ListSteps`, `ListInstanceGroups`, `ListInstanceFleets`, `ListInstances`, `ListBootstrapActions`, `GetAutoTerminationPolicy` | Discovers clusters and steps, sizes them for cost, and reads scaling and termination settings for recommendations.                                                                                                                                                                           | Change any setting.                                                          |
| `elasticmapreduce:CreatePersistentAppUI`, `DescribePersistentAppUI`, `GetPersistentAppUIPresignedURL`                                                                            | Opens the Persistent App UI, the off-cluster Spark History Server that EMR keeps for 30 days after a cluster ends, and reads the Spark event log from it. AWS classifies the create action as a write because it starts a UI session; the session exposes only that cluster's Spark history. | Read S3, modify clusters, or reach anything outside that cluster's Spark UI. |
| S3                                                                                                                                                                               | None. Event logs are read through the Persistent App UI.                                                                                                                                                                                                                                     |                                                                              |

{% hint style="warning" %}
**Prerequisite:** the cluster keeps Spark's default event-log directory in HDFS. If `spark.eventLog.dir` is redirected to S3, EMR does not build the Persistent App UI for that cluster and its runs are not ingested.
{% endhint %}

### Policy

Policy name: `DataflintEmrReadOnly`. Role name: `dataflint-emr-read-only-role`.

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DataflintEmrEc2ReadOnly",
      "Effect": "Allow",
      "Action": [
        "elasticmapreduce:ListClusters",
        "elasticmapreduce:DescribeCluster",
        "elasticmapreduce:ListSteps",
        "elasticmapreduce:ListInstanceGroups",
        "elasticmapreduce:ListInstanceFleets",
        "elasticmapreduce:ListInstances",
        "elasticmapreduce:ListBootstrapActions",
        "elasticmapreduce:GetAutoTerminationPolicy",
        "elasticmapreduce:CreatePersistentAppUI",
        "elasticmapreduce:DescribePersistentAppUI",
        "elasticmapreduce:GetPersistentAppUIPresignedURL"
      ],
      "Resource": "*"
    }
  ]
}
```

{% endtab %}

{% tab title="EMR Serverless" %}

### Why each permission is needed

| Actions                                                                              | What DataFlint does with it                                                                                                                     | What it cannot do                                            |
| ------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ |
| `emr-serverless:ListApplications`                                                    | Discovers your Spark applications.                                                                                                              | Create or change applications.                               |
| `emr-serverless:GetApplication`, `ListJobRuns`, `GetJobRun`, `GetDashboardForJobRun` | Reads application and job-run configuration and status.                                                                                         | Submit, cancel or change job runs.                           |
| `s3:ListBucket`, `s3:GetObject` on your log bucket                                   | Reads the Spark event log and driver log that EMR Serverless writes to the S3 location in your application's or job's monitoring configuration. | Write or delete, or read outside the prefix you scope it to. |

{% hint style="warning" %}
**Prerequisite:** S3 logging must be enabled on the application or on each job (`monitoringConfiguration.s3MonitoringConfiguration.logUri`). Runs without S3 logging are not ingested.
{% endhint %}

### Policy

Policy name: `DataflintEmrServerlessReadOnly`. Role name: `dataflint-emr-serverless-read-only-role`. Replace `YOUR_REGION`, `YOUR_ACCOUNT_ID` and `YOUR_LOG_BUCKET`.

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowEmrServerlessList",
      "Effect": "Allow",
      "Action": "emr-serverless:ListApplications",
      "Resource": "*"
    },
    {
      "Sid": "AllowEmrServerlessJobReads",
      "Effect": "Allow",
      "Action": [
        "emr-serverless:GetApplication",
        "emr-serverless:ListJobRuns",
        "emr-serverless:GetJobRun",
        "emr-serverless:GetDashboardForJobRun"
      ],
      "Resource": [
        "arn:aws:emr-serverless:YOUR_REGION:YOUR_ACCOUNT_ID:/applications/*",
        "arn:aws:emr-serverless:YOUR_REGION:YOUR_ACCOUNT_ID:/applications/*/jobruns/*"
      ]
    },
    {
      "Sid": "AllowSparkEventLogRead",
      "Effect": "Allow",
      "Action": [
        "s3:ListBucket",
        "s3:GetObject"
      ],
      "Resource": [
        "arn:aws:s3:::YOUR_LOG_BUCKET",
        "arn:aws:s3:::YOUR_LOG_BUCKET/*"
      ]
    }
  ]
}
```

{% endtab %}

{% tab title="EMR on EKS" %}

### Why each permission is needed

| Actions                                                                                               | What DataFlint does with it                                                                                                 | What it cannot do                                                                |
| ----------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |
| `emr-containers:ListVirtualClusters`, `DescribeVirtualCluster`, `ListJobRuns`, `DescribeJobRun`       | Discovers virtual clusters and job runs and reads their configuration.                                                      | Submit or cancel job runs.                                                       |
| `elasticmapreduce:CreatePersistentAppUI`, `DescribePersistentAppUI`, `GetPersistentAppUIPresignedURL` | Opens the Persistent App UI for a job run when its event log is not available in S3.                                        | Modify anything.                                                                 |
| `eks:DescribeCluster`, `ListNodegroups`, `DescribeNodegroup`                                          | Reads node-group instance types so cost is attributed correctly.                                                            | Read pods, secrets or any Kubernetes object. The Kubernetes API is never called. |
| `s3:ListBucket`, `s3:GetObject` on your log bucket                                                    | Reads Spark event logs from `spark.eventLog.dir` or your job's S3 monitoring location and, for failed runs, the driver log. | Write or delete, or read outside the prefix you scope it to.                     |

{% hint style="warning" %}
**Prerequisite:** job runs write Spark event logs to S3, either through `spark.eventLog.dir` in the job's Spark configuration or through `monitoringConfiguration.s3MonitoringConfiguration.logUri`.
{% endhint %}

### Policy

Policy name: `DataflintEmrContainersReadOnly`. Role name: `dataflint-emr-eks-read-only-role`. Replace `YOUR_LOG_BUCKET`.

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowEmrContainersReadOnly",
      "Effect": "Allow",
      "Action": [
        "emr-containers:ListVirtualClusters",
        "emr-containers:DescribeVirtualCluster",
        "emr-containers:ListJobRuns",
        "emr-containers:DescribeJobRun",
        "elasticmapreduce:CreatePersistentAppUI",
        "elasticmapreduce:DescribePersistentAppUI",
        "elasticmapreduce:GetPersistentAppUIPresignedURL",
        "eks:DescribeCluster",
        "eks:ListNodegroups",
        "eks:DescribeNodegroup"
      ],
      "Resource": "*"
    },
    {
      "Sid": "AllowSparkEventLogRead",
      "Effect": "Allow",
      "Action": [
        "s3:ListBucket",
        "s3:GetObject"
      ],
      "Resource": [
        "arn:aws:s3:::YOUR_LOG_BUCKET",
        "arn:aws:s3:::YOUR_LOG_BUCKET/*"
      ]
    }
  ]
}
```

{% endtab %}
{% endtabs %}

## Create the role

Pick one method. All methods create the same role: a permissions policy from the tab above, attached to a role that trusts the DataFlint principal with your external ID.

{% tabs %}
{% tab title="AWS Console (UI)" %}

### Step 1: Create the IAM policy

1. Open **IAM → Policies → Create policy**.
2. Choose **JSON** and paste the policy from your EMR service's tab.
3. Name it as shown in the tab (for example `DataflintEmrReadOnly`) and create it.

### Step 2: Create the IAM role

1. Open **IAM → Roles → Create role**.
2. Trusted entity type: **Custom trust policy**. Paste the trust policy from this page, with your external ID.
3. Attach the policy from Step 1.
4. Name the role as shown in the tab and create it.

### Step 3: Copy the role ARN

Open the role and copy its **ARN**. You will paste it into the DataFlint console.
{% endtab %}

{% tab title="AWS CLI" %}
Save the trust policy as `trust-policy.json` and the policy from your tab as `permissions-policy.json`, then run (example names are for EMR on EC2):

```bash
ROLE_NAME="dataflint-emr-read-only-role"
POLICY_NAME="DataflintEmrReadOnly"

aws iam create-role \
  --role-name "${ROLE_NAME}" \
  --assume-role-policy-document file://trust-policy.json \
  --description "Read-only role for DataFlint EMR ingestion"

POLICY_ARN="$(aws iam create-policy \
  --policy-name "${POLICY_NAME}" \
  --policy-document file://permissions-policy.json \
  --query 'Policy.Arn' \
  --output text)"

aws iam attach-role-policy \
  --role-name "${ROLE_NAME}" \
  --policy-arn "${POLICY_ARN}"

aws iam get-role --role-name "${ROLE_NAME}" --query 'Role.Arn' --output text
```

{% endtab %}

{% tab title="Terraform" %}
Save the policy from your tab as `permissions-policy.json` next to this file.

```hcl
variable "dataflint_external_id" {
  description = "External ID shown in the DataFlint console"
  type        = string
  sensitive   = true
}

variable "role_name" {
  type    = string
  default = "dataflint-emr-read-only-role"
}

resource "aws_iam_role" "dataflint_emr_read_only" {
  name = var.role_name

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect    = "Allow"
      Principal = { AWS = "arn:aws:iam::975050001706:role/eks-dataflint-service-role" }
      Action    = "sts:AssumeRole"
      Condition = { StringEquals = { "sts:ExternalId" = var.dataflint_external_id } }
    }]
  })
}

resource "aws_iam_policy" "dataflint_emr_read_only" {
  name   = "${var.role_name}-policy"
  policy = file("${path.module}/permissions-policy.json")
}

resource "aws_iam_role_policy_attachment" "dataflint_emr_read_only" {
  role       = aws_iam_role.dataflint_emr_read_only.name
  policy_arn = aws_iam_policy.dataflint_emr_read_only.arn
}

output "role_arn" {
  value = aws_iam_role.dataflint_emr_read_only.arn
}
```

```bash
terraform init
terraform apply
terraform output role_arn
```

{% endtab %}

{% tab title="CloudFormation" %}
The DataFlint console generates a CloudFormation template for your EMR service with the policy, the principal and your external ID filled in (Admin Panel → Environment Management → Add New → EMR → "Need to create this role?" → CloudFormation). Deploy it with:

```bash
aws cloudformation deploy \
  --template-file dataflint-emr-read-role.yaml \
  --stack-name dataflint-emr-read-role \
  --capabilities CAPABILITY_NAMED_IAM \
  --region YOUR_REGION
```

Compare the template's policy statement with the tab above before deploying.
{% endtab %}
{% endtabs %}

## Add it in DataFlint

In the DataFlint console go to **Admin Panel → Environment Management → Add New → EMR** and fill in:

* **Environment name**
* **Region** — one region per environment. Add another environment for each additional region.
* **Engine type** — EC2, EKS or Serverless.
* **IAM role ARN** — the role you created above.
* **Cluster ID** (optional) — restrict ingestion to one cluster or virtual cluster.

Click **Add Pipeline**. The environment shows as **Pending approval** while DataFlint verifies access; ingestion starts once it is approved.

## Scope it down

* **Region:** add `"Condition": {"StringEquals": {"aws:RequestedRegion": "YOUR_REGION"}}` to the EMR statement.
* **Buckets:** the S3 statement already names one bucket. Restrict listing to the log prefix with `"Condition": {"StringLike": {"s3:prefix": ["YOUR_PREFIX/*"]}}` on `s3:ListBucket`.
* **Encrypted logs:** if the log bucket uses SSE-KMS, also allow `kms:Decrypt` on the key.

## Validate the setup

```bash
ROLE_ARN="$(aws iam get-role --role-name dataflint-emr-read-only-role --query 'Role.Arn' --output text)"

# EMR on EC2
aws iam simulate-principal-policy --policy-source-arn "${ROLE_ARN}" \
  --action-names elasticmapreduce:ListClusters elasticmapreduce:CreatePersistentAppUI elasticmapreduce:GetPersistentAppUIPresignedURL \
  --output text

# EMR Serverless
aws iam simulate-principal-policy --policy-source-arn "${ROLE_ARN}" \
  --action-names emr-serverless:ListApplications emr-serverless:GetJobRun s3:GetObject \
  --output text

# EMR on EKS
aws iam simulate-principal-policy --policy-source-arn "${ROLE_ARN}" \
  --action-names emr-containers:ListVirtualClusters emr-containers:DescribeJobRun eks:DescribeCluster s3:GetObject \
  --output text
```

Every line should show `allowed`.
