S3
Enables seamless streaming of data posted to an S3 bucket into the Monad solution.
Sync Type: Incremental
Requirements
- IAM Role Assumption / Static Credentials
- Example permission to attach to the role/user:
Code
Details
By default, the first run ingests only files written from the time the input starts onward — no historical objects are fetched. To ingest history, set a Backfill Start Time: the first run then performs a one-time sync of all files in the specified bucket-prefix from that date up to now.
On subsequent runs, the processor performs an incremental sync starting from last checkpoint timestamp it found a blob at. On restart in case of any form of failure, we resume from the day prefix of the last checkpointed timestamp. A checkpoint occurs at every page within a prefix. So while processing a prefix, if a failure occurs, the processor will restart from the last completed page's checkpoint. You will not lose out on any records, however you may re-process some data in the S3 objects on the page where the failure occured in case of any catastrophic failures.
-
To avoid this, we recommend publishing S3 data to an SQS queue to avoid such failures.
-
Please also note we rescan and drop all data based on our deduplication logic on every single sync which occurs in a day prefix. This means that for larger buckets, this may lead to hitting rate limits since we will be scanning the same data a large number of times in a day. To avoid this, we recommend publishing S3 data to an SQS queue to avoid such scenarios.
-
The date partition must directly follow the Prefix and must use one of the supported Partition Formats (simple date, hive compliant, flat hive compliant, or compact date). Anything other than this can cause unexpected behavior in the input.
-
Each log's last updated time should be on the same date as the logical prefix itself. so any object that lands in the 2025/08/10 prefix should have a last updated time of 2025/08/10 (in its ISO8601 format). Not doing so can cause unexpected behavior in the input.
-
To avoid such tight boundaries, we recommend publishing S3 data to an SQS queue to avoid such failures.
The processor checks for new data continuously every 10 seconds, processing objects as they appear in the bucket.
Configuration
The following configuration defines the input parameters. Each field's specifications, such as type, requirements, and descriptions, are detailed below.
Settings
| Setting | Type | Required | Description |
|---|---|---|---|
| Region | string | No | The region of the S3 bucket. If left blank, the region will be auto-detected. |
| Bucket | string | Yes | The name of the S3 bucket. |
| Prefix | string | No | Prefix of the S3 object keys to read, up to (but not including) the date partition — e.g. AWSLogs/123456789012/CloudTrail/us-west-2. See Prefix below. |
| Key Filter | object | No | Optional filter applied to the S3 object key before fetch. See Key Filter below. |
| Compression | string | Yes | Compression format of the S3 objects. |
| Format | string | Yes | File format of the S3 objects. |
| Partition Format | string | Yes | The existing partition format used in your S3 bucket. |
| RoleARN | string | Yes | Role ARN to assume when reading from S3. |
| Record Location | string | No | Location of the record in the JSON object. See Record Location for syntax and examples. |
| Schema | array | No | Ordered list of column names for headerless delimited files. Applies to the delimited format only. See Schema below. |
| Backfill Start Time | string | No | The date to start fetching data from. If not specified, no past records will be fetched. |
Prefix
The Prefix is the literal key path that sits directly in front of the date partition in your bucket. The connector appends the date partition to it, so the Prefix must stop immediately before the date. Leave it blank when the date partitions sit at the root of the bucket.
Write a multi-level prefix as a plain /-separated path.
Leading and trailing slashes are ignored, so /logs/app/ and logs/app are equivalent.
Wildcards and placeholders are not supported.
| Prefix | Partition Format | Object keys the connector reads |
|---|---|---|
| (blank) | simple date | 2026/09/11/events.json.gz |
logs | simple date | logs/2026/09/11/events.json.gz |
AWSLogs/123456789012/CloudTrail/us-west-2 | simple date | AWSLogs/123456789012/CloudTrail/us-west-2/2026/09/11/events.json.gz |
exports/audit | flat hive compliant | exports/audit/dt=2026-09-11/events.json.gz |
lake/findings | compact date | lake/findings/eventDay=20260911/events.json.gz |
Keep in mind:
- S3 matches the Prefix as a literal string, so
logs/appalso matches keys underlogs/app-old/. Use a more specific prefix, or a Key Filter, to exclude them. - One input reads one prefix. If the same layout repeats under several prefixes (for example one per account or region), create one S3 input per prefix. For CloudTrail logs, the CloudTrail input discovers account and region prefixes on its own.
Key Filter
Filters S3 object keys before the connector fetches them. Useful when a single prefix holds multiple object types distinguished by a token in the file name (e.g. Amazon Redshift audit logs, where connectionlog and useractivitylog objects share one prefix), letting each type be split across separate pipelines. Matching is case-insensitive on both the key and the configured value.
Key Filter is a discriminated union on the type field — the empty case is {"type": "none"} rather than an empty {mode, operator, value} object, so the per-field rules only apply when filtering is actually requested:
| Field | Type | Required | Description |
|---|---|---|---|
type | string | Yes | none passes every key through; filter enables the rules in filter below |
filter | object | When type is filter | The match parameters described below. Omit when type is none. |
When type is filter, the nested filter object takes:
| Field | Type | Required | Description |
|---|---|---|---|
mode | string | Yes | include processes only matching keys; exclude skips matching keys |
operator | string | Yes | One of contains, ends_with, starts_with |
value | string | Yes | The substring (or suffix/prefix) to match against the S3 object key |
Examples:
- No filtering —
{ "type": "none" } - Process only the connection logs —
{ "type": "filter", "filter": { "mode": "include", "operator": "contains", "value": "connectionlog" } } - Skip a noisy prefix —
{ "type": "filter", "filter": { "mode": "exclude", "operator": "starts_with", "value": "internal/" } }
When the field is omitted entirely (or set to {"type": "none"}), every S3 object under the prefix is processed.
Schema
An ordered list of column names for headerless delimited files (e.g. PSV), such as Amazon Redshift audit logs.
This setting applies to the delimited format only. The csv and wsv formats always read column names from the first row and ignore this field. When the schema is left empty, the delimited format also falls back to reading the header from the first row.
Example — supplying column names for a headerless pipe-separated file:
Code
Partition Format
The Partition Format setting specifies the existing organization of data within your S3 bucket. This is crucial for the system to correctly navigate and read your data. Select the option that matches your current S3 bucket structure:
-
Simple Date Format ('simple date'):
- Structure:
YYYY/MM/DD - Example:
2024/01/01 - Use case: For buckets using basic chronological organization of data.
- Structure:
-
Hive-compatible Format ('hive compliant'):
- Structure:
year=YYYY/month=MM/day=DD - Example:
year=2024/month=01/day=01 - Use case: For buckets set up in a Hive-compatible format, common in data lake configurations.
- Structure:
-
Flat Hive Compliant Format ('flat hive compliant'):
- Structure:
dt=YYYY-MM-DD - Example:
dt=2024-01-01 - Use case: For buckets using a single-key hive-style date partition.
- Structure:
-
Compact Date Format ('compact date'):
- Structure:
eventDay=YYYYMMDD - Example:
eventDay=20240101 - Use case: For buckets that write a single
eventDay=key with an undelimited date, such as Amazon Security Lake exports.
- Structure:
Selecting the correct Partition Format ensures that the system can efficiently locate and process your existing data by matching your S3 bucket's current structure. This setting does not change your bucket's organization; it tells the system how to navigate it.
Secrets (Static Credentials Only)
| Setting | Type | Required | Description |
|---|---|---|---|
| Access Key | string | Conditional | AWS Access Key ID |
| Secret Key | string | Conditional | AWS Secret Access Key |
⚠️ Authentication: Choose either Role ARN (recommended) or static credentials. See AWS Authentication Guide for setup instructions.
Sync frequency
This input polls on a connector-specific interval, adjusting between about 10 seconds and 5 minutes depending on the backlog. A cron schedule configured on the pipeline overrides this cadence. See Input Sync Frequency for details.