Cheat Sheets/DEA-C01
AssociateDEA-C01

AWS Certified Data Engineer – Associate Cheat Sheet (DEA-C01)

Every in-scope ingestion, transformation, storage and governance service the DEA-C01 exam actually tests — organised by the four exam domains with decision tables and runbooks

Free PDF — no signup required

13 pages1.0 MBUpdated April 24, 2026

About This Cheat Sheet

A 13-page associate-tier reference covering every in-scope service from the official DEA-C01 exam guide, organised around the four exam domains: Data Ingestion & Transformation (34%), Data Store Management (26%), Data Operations & Support (22%), Data Security & Governance (18%). Structured for candidates with ~2 years of data-engineering experience — sections move from fundamentals (structured / semi-structured / unstructured data, OLTP vs OLAP, batch vs streaming, lake vs warehouse vs lakehouse, medallion architecture, ETL vs ELT, partitioning vs bucketing, columnar formats Parquet / ORC / Avro) through ingestion (Kinesis Data Streams with shard limits + Enhanced Fan-Out + On-Demand, Kinesis Data Firehose with Lambda transform + Parquet/ORC conversion + dynamic partitioning, MSK + MSK Serverless + MSK Connect, Managed Service for Apache Flink, DataSync, AWS DMS with full-load + CDC, Transfer Family, Application Migration / Discovery Service, Snow Family for PB-scale offline, AppFlow for SaaS connectors), transformation + orchestration (Glue with Data Catalog + Crawlers + Jobs + Workflows + Triggers + Bookmarks + DataBrew + Data Quality DQDL + Schema Registry + Sensitive Data Detection + Glue Studio + Interactive Sessions, EMR Serverless + EMR on EC2 + EMR on EKS with Spark / Hive / Presto / Trino / HBase / Flink + table formats Iceberg / Hudi / Delta Lake, Lambda, AWS Batch, Amazon ECS / EKS / ECR for custom ETL workers, Step Functions with distributed Map, Amazon MWAA for Airflow-native DAGs, EventBridge + Scheduler, SNS / SQS), storage (S3 with storage classes + Intelligent-Tiering + Lifecycle + Versioning + Object Lock + CRR + Replication Time Control + Access Points + S3 Tables managed Iceberg + S3 Glacier + S3 Select deprecation, EBS, EFS, AWS Backup), analytical engines (Athena with workgroups + Federated Query + Athena for Spark + CTAS / UNLOAD, Redshift with RA3 + Serverless + DISTKEY / SORTKEY / ENCODE + Spectrum + Zero-ETL + data sharing + Redshift ML + Streaming Ingestion + Concurrency Scaling + Auto WLM + materialized views, EMR engine reference, OpenSearch with UltraWarm + Cold tiers + k-NN + Neural plugin), data catalog + lakehouse + governance (Glue Data Catalog with partition projection, AWS Lake Formation FGAC with LF-Tags + row/column/cell-level, AWS Data Exchange), databases (RDS, Aurora, DynamoDB + Streams + Global Tables, DocumentDB, Keyspaces, MemoryDB, Neptune, Redshift), visualisation (Amazon QuickSuite — 2025 rebrand unifying QuickSight + Q with SPICE + Q for NL BI + row-level security + embedded analytics, Managed Grafana, SageMaker AI for ML-in-pipeline, Bedrock for enrichment, Kendra, Amazon Q), streaming patterns (Lambda vs Kappa architectures, CDC with Zero-ETL, exactly-once semantics), security (IAM, KMS, Secrets Manager, Macie, WAF / Shield, CloudTrail, Config, at-rest + in-transit + client-side encryption, Lake Formation FGAC), networking (VPC + gateway/interface endpoints, PrivateLink, Route 53, CloudFront, API Gateway), observability (CloudWatch + CloudWatch Logs, Managed Grafana, Systems Manager, Well-Architected Tool Data Analytics + ML Lenses, Budgets + Cost Explorer), DevOps (CloudFormation, CDK, SAM, CLI, CodeBuild / CodeDeploy / CodePipeline, Amazon Q Developer), five reference architectures (serverless lakehouse, streaming ETL with CDC, enterprise data mesh, event-driven batch pipeline, hybrid ingest for legacy), 22-scenario answer patterns, 12 common pitfalls, Glue deep dive (job types, bookmarks, Data Quality DQDL), Redshift performance playbook (physical design, ingest patterns, concurrency + sharing), lakehouse table formats (Iceberg / Hudi / Delta), four operational runbooks (Glue failure, Kinesis lag, Redshift slow query, Athena cost spike), data quality + lineage + compliance framework, and a 17-row architectural decision cheatsheet. Every service on the DEA-C01 in-scope list is audited present.

What's Inside

1

Data Engineering Fundamentals

Structured / semi-structured / unstructured data; OLTP vs OLAP; batch vs streaming; data lake vs warehouse vs lakehouse; medallion architecture (Bronze / Silver / Gold); ETL vs ELT; partitioning vs bucketing; columnar formats (Parquet as the default, ORC for Hive ecosystems, Avro for streaming with Schema Registry).

2

Ingestion

Amazon Kinesis Data Streams (shard limits 1 MB/s in + 2 MB/s out, Enhanced Fan-Out, On-Demand), Kinesis Data Firehose (managed delivery with Lambda transform + Parquet/ORC conversion + dynamic partitioning), Amazon MSK + MSK Serverless + MSK Connect, Amazon Managed Service for Apache Flink (Kinesis Data Analytics successor — SQL + DataStream + CEP + windows), AWS DataSync, AWS DMS (full-load + CDC, serverless, SCT for schema conversion), AWS Transfer Family (SFTP/FTPS/FTP/AS2/WebApp), AWS Application Discovery + Migration Service, AWS Snow Family for PB-scale offline, Amazon AppFlow SaaS connectors.

3

Transformation & Orchestration

AWS Glue (Data Catalog, Crawlers, Jobs including Spark + Spark Streaming + Python Shell + Ray, Workflows + Triggers, Bookmarks, DataBrew low-code, Data Quality with DQDL, Schema Registry, Sensitive Data Detection, Glue Studio, Interactive Sessions), Amazon EMR (Serverless + EC2 + EKS with Spark / Hive / Presto / Trino / HBase / Flink and table formats Iceberg / Hudi / Delta), AWS Lambda, AWS Batch, Amazon ECS / EKS / ECR for custom ETL, AWS Step Functions (Standard + Express + distributed Map), Amazon MWAA for Airflow-native DAGs, Amazon EventBridge + Scheduler, Amazon SNS + SQS (including FIFO for ordering + extended client for >256 KB payloads).

4

Storage

Amazon S3 analytical lake: storage classes (Standard, Intelligent-Tiering, IA, One Zone-IA, Glacier Instant / Flexible / Deep Archive), Lifecycle policies, Versioning + Object Lock, Replication (SRR / CRR + RTC 15-min SLA), Access Points + Multi-Region Access Points, Event Notifications for reactive pipelines. Amazon S3 Tables for managed Apache Iceberg with auto-compaction + snapshot management. Amazon S3 Glacier with Vault Lock for compliance. Amazon EBS gp3 / io2 Block Express, Amazon EFS, AWS Backup for org-wide backup plans across EBS / EFS / RDS / Aurora / DynamoDB / FSx / S3.

5

Analytical Engines & Warehouses

Amazon Athena (Presto / Trino SQL on S3, workgroups with per-query scan limit + enforced encryption, Federated Query via Lambda connectors, Athena for Apache Spark, CTAS / UNLOAD for partitioned output). Amazon Redshift (RA3 managed storage, DC2 fixed, Redshift Serverless pay-per-RPU-second, DISTKEY / SORTKEY / ENCODE / VACUUM / ANALYZE, Spectrum over S3 including Iceberg, Zero-ETL from Aurora / RDS / DynamoDB / OpenSearch, data sharing across cluster / account / Region, Redshift ML, Streaming Ingestion, Concurrency Scaling, materialized views + Auto Materialized Views, Auto WLM with query monitoring rules). Amazon OpenSearch Service (UltraWarm + Cold, k-NN + Neural plugin, SAML SSO, domains + Serverless).

6

Catalog, Lakehouse & Governance

AWS Glue Data Catalog as central metastore used by Athena, EMR, Redshift Spectrum, Lake Formation, SageMaker; Crawlers for schema inference; partition projection in Athena to avoid catalog partition explosion. AWS Lake Formation fine-grained access control (row / column / cell) with LF-Tags, cross-account + cross-Region sharing, hybrid-mode migration from IAM-only. AWS Data Exchange for subscribing to third-party data products as S3 datasets, Redshift shares, API, or LF grants.

7

Databases

Amazon RDS (MySQL / PostgreSQL / MariaDB / Oracle / SQL Server / Db2 with Multi-AZ + read replicas + Performance Insights + Blue/Green deployments), Amazon Aurora (compatible MySQL / PostgreSQL with 128 TiB storage auto-growth + 15 read replicas + Global Database + Serverless v2 + Zero-ETL to Redshift), Amazon DynamoDB (+ Streams, GSI / LSI, TTL, Transactions, PITR, Global Tables active-active, DAX), Amazon DocumentDB (MongoDB-compatible), Amazon Keyspaces (Cassandra-compatible serverless), Amazon MemoryDB for Redis (durable, Multi-AZ transaction log), Amazon Neptune (graph + Neptune Analytics vector + graph), Amazon Redshift.

8

Streaming Patterns

Lambda architecture (batch + speed merged view), Kappa architecture (single streaming path), Change Data Capture with DMS → Kinesis or Zero-ETL, event-driven lake (S3 PUT → EventBridge → Step Functions → Glue → Athena), backfill + reconciliation with Iceberg MERGE. Exactly-once + ordering: Kinesis at-least-once + idempotent downstream writes (Iceberg MERGE, DynamoDB conditional PutItem); Kafka idempotent producer + transactions for exactly-once within one app; SQS FIFO 3,000 msg/s with message-group exactly-once.

9

Security, Networking & Compliance

IAM least-privilege roles for Glue / EMR / Lambda / SageMaker; KMS customer-managed keys everywhere; Secrets Manager with rotation for DB + third-party API keys; Macie for PII in S3; WAF + Shield on data-intake APIs; CloudTrail org trail for every Glue / Athena / Redshift Data API call; Config managed rules + auto-remediation. Encryption: at-rest with KMS, in-transit with TLS (rds.force_ssl, Redshift require_ssl), client-side with AWS Encryption SDK, column / row masking via Lake Formation and Redshift dynamic data masking. Networking: VPC with private subnets for data stores, free S3 / DynamoDB Gateway endpoints, PrivateLink interface endpoints for Glue / Athena / Kinesis / MSK / KMS / Secrets Manager / SageMaker, Route 53, CloudFront, API Gateway with WAF + authorizers.

10

Observability, Reliability & Cost

CloudWatch metrics (Glue bytesRead, Kinesis IteratorAgeMilliseconds, Redshift WLMQueueLength, RDS CPUUtilization), alarms (standard + composite + anomaly + metric math), dashboards + Synthetics; CloudWatch Logs with Logs Insights + subscription filters → Firehose → OpenSearch; Amazon Managed Grafana for unified dashboards; AWS Systems Manager for Parameter Store + Patch Manager + Change Manager + Incident Manager; AWS Well-Architected Tool with Data Analytics + Machine Learning Lenses; AWS Budgets with actions + Cost Explorer. Cost patterns: Parquet + partitioning + compression for Athena, RA3 + Spectrum + Concurrency Scaling + Reserved Nodes for Redshift, EMR Serverless or Spot task nodes, S3 Intelligent-Tiering + Lifecycle, Kinesis On-Demand, Glue Flex + Auto Scaling.

11

Reference Architectures

Five full walkthroughs: serverless lakehouse (Firehose → S3 Parquet → Glue PySpark → Iceberg via S3 Tables → Athena / Redshift Spectrum / QuickSuite), streaming ETL with CDC (Aurora Zero-ETL to Redshift or DMS CDC → Kinesis → Managed Flink → S3 Iceberg with DLQ + backfill), enterprise data mesh (domain accounts own data + central governance + Lake Formation cross-account grants + discovery via SageMaker Unified Studio / QuickSuite / Q Business), event-driven batch pipeline (S3 PUT → EventBridge → Step Functions → Glue Data Quality gate → Athena + SNS alerts), hybrid ingest for legacy systems (DataSync + Transfer Family + DMS + Snow Family + Application Migration Service all landing in governed S3 / Lake Formation).

12

Scenarios, Pitfalls & Runbooks

Twenty-two scenario → answer mappings (cheap serverless SQL, FGAC on lake, near-real-time Aurora→Redshift, schema evolution for streaming, ACID on lake, PII classification, Airflow, replace on-prem Hadoop, SaaS ingest, deduplication, cross-account Redshift, federated query, rotating secrets, late-arriving records, account isolation, Kinesis Data Analytics replacement, batch containers, Athena cost, cross-Region warehouse, natural-language BI, reprocessing, lineage); twelve common pitfalls (Firehose vs Streams, Crawler catalog explosion, Athena scan cost, Redshift WLM, DMS sizing, EMR Spot placement, Kinesis limits, LF vs IAM, S3 Tables vs DIY Iceberg, Zero-ETL latency, data sovereignty, Athena cost explosion); four runbooks (Glue job failure, Kinesis consumer lag, Redshift slow query, Athena cost spike).

Why This Cheat Sheet Helps

DEA-C01 is breadth-heavy: you must pick the right service for each stage of a pipeline — ingest, transform, store, secure, observe — without wasting money or introducing latency. Questions reliably distinguish Firehose from Streams, Glue from EMR, Athena partitioning from Redshift dist-keys, Lake Formation grants from S3 bucket policies, Zero-ETL from DMS CDC, S3 Tables (managed Iceberg) from classic Iceberg on S3. This cheat sheet puts those decisions side by side with the exact AWS vocabulary examiners use.

It assumes you've already written some Spark / SQL / Python on AWS. Use it in the final weeks to reinforce the Glue + Lake Formation + Athena + Redshift core, rehearse the streaming patterns (Kinesis / MSK / Flink / Zero-ETL), and memorise the operational runbooks for the failures the exam loves (Glue OOM / skew, Kinesis consumer lag, Redshift slow query, Athena cost spike).

How to Use It

Skim the whole sheet once to see how the four domains map to sections. Then hammer practice exams — for each wrong answer, find the relevant section and study the tables and decision trees. Pay extra attention to sections 2 (Ingestion), 3 (Transformation + Orchestration), 5 (Analytical engines), 6 (Catalog + Lake Formation), 15 (Scenario → Answer Patterns), 16 (Pitfalls), 17 (Glue deep dive), 18 (Redshift performance), and 22 (Architectural decision cheatsheet).

In the final week, walk through the four domain weights against your confidence: Ingestion & Transformation (34%) → sections 2–3, 9; Data Store Management (26%) → sections 4–7, 19; Operations & Support (22%) → sections 12–13, 20; Security & Governance (18%) → sections 6, 10, 21. Pair with CloudNinja's free DEA-C01 practice exam to surface gaps.

Frequently Asked Questions

Is this AWS Data Engineer Associate cheat sheet free?

Yes, completely free with no signup required. Download the PDF directly from CloudNinja and use it as a study reference for the DEA-C01 exam.

How much experience do I need before DEA-C01?

AWS recommends 2–3 years of data-engineering experience (ideally on AWS) and 1–2 years of hands-on AWS. You should be comfortable writing SQL and Python / PySpark, using S3 + Glue + Athena, and understanding at least one streaming technology (Kinesis or Kafka). If you do not already hold Solutions Architect Associate or have production data-pipeline experience on AWS, build that foundation first.

How is DEA-C01 different from the Solutions Architect Associate?

SAA-C03 spans all AWS services and tests which service fits a use case; DEA-C01 goes deep on the data lifecycle. Expect detailed decisions: Glue vs EMR vs Lambda for a given transform, Firehose vs Streams for an ingest pattern, Lake Formation row-level vs column-level grants, Iceberg vs Hudi, Redshift DISTKEY vs SORTKEY choices, Zero-ETL vs DMS CDC, S3 Tables vs classic Iceberg. Questions are scenario-heavy and lean on tradeoffs around cost, latency, and operational overhead.

Is this cheat sheet updated for the current DEA-C01 exam?

Yes, it is built directly from the current DEA-C01 exam guide (data-engineer-associate-01) and audited against the full in-scope services list. It reflects current services and recent changes — Amazon S3 Tables (managed Apache Iceberg), Amazon QuickSuite (the 2025 unification of QuickSight + Amazon Q), Zero-ETL integrations (Aurora / RDS MySQL + PostgreSQL / DynamoDB / OpenSearch to Redshift), Amazon Managed Service for Apache Flink as the successor to Kinesis Data Analytics, Redshift Streaming Ingestion, Glue Data Quality DQDL, Amazon Q Developer, and the S3 Select deprecation for new customers.

Keep Studying

We use cookies to improve your experience. This site uses YouTube embeds and Google Analytics to understand how visitors use our site. Learn more