AWS
Cloud Asset Inventory
Product News
AWS Relations Transformer: Every Dependency in One Table
If you sync AWS with CloudQuery, the data to answer "what is connected to what" is already in your database. Getting it out is the hard part: an EC2 instance points at its VPC through
vpc_id, at its IAM role through iam_instance_profile_arn, and at its security groups through a nested list. Each service does it differently, across hundreds of tables.Why Did We Build the AWS Relations Transformer? #
Most teams we talk to write the same joins over and over, one relationship at a time. We wanted that work done once, during the sync. We also expect customers to find uses we haven't thought of, which is why we're shipping this as a preview and asking for input.
How Does the AWS Relations Transformer Work? #
It reads the AWS records flowing into your destination and writes one row per reference to
aws_resource_relationships. We call each row an edge. Every edge has the _cq_id of both resources, the column the reference came from, and the ARN or ID on each side:select source_identifier, source_column, target_table
from aws_resource_relationships
where target_identifier = 'arn:aws:ec2:us-east-1:123456789012:vpc/vpc-0abc';
Each edge keeps the same
_cq_id across syncs, so you can track how relationships change over time.Can I Walk More Than One Hop? #
You can, but walk the edges in both directions. AWS records containment from child to parent: a subnet has a
vpc_id, while a VPC has nothing about its subnets. In one of our test accounts, an outbound-only walk from a VPC found a single DHCP option set. The same walk in both directions found 74 resources.This PostgreSQL view lists every edge in both directions. It leaves out account and region edges, which would otherwise put every resource in an account two hops from every other one:
create view aws_resource_relationship_adjacency as
select source_table as a_table, source_cq_id as a_cq_id,
target_table as b_table, target_cq_id as b_cq_id
from aws_resource_relationships
where relationship_type not in ('account_reference', 'region_reference')
union
select target_table, target_cq_id, source_table, source_cq_id
from aws_resource_relationships
where relationship_type not in ('account_reference', 'region_reference');
With the view in place, this query returns every resource within three hops of a VPC, with the fewest hops it takes to reach each one:
with recursive walk as (
select 'aws_ec2_vpcs'::text as tbl, _cq_id as id, 0 as depth,
array['aws_ec2_vpcs:' || _cq_id::text] as path
from aws_ec2_vpcs
where arn = 'arn:aws:ec2:us-east-1:111122223333:vpc/vpc-0abc'
union all
select e.b_table, e.b_cq_id, w.depth + 1,
w.path || (e.b_table || ':' || e.b_cq_id::text)
from walk w
join aws_resource_relationship_adjacency e
on e.a_table = w.tbl and e.a_cq_id = w.id
where w.depth < 3
and not (e.b_table || ':' || e.b_cq_id::text) = any(w.path)
)
select tbl, id, min(depth) as hop
from walk group by 1, 2 order by hop, tbl;
How Do I Set It Up? #
One thing to know first: a transformer replaces the records sent to the destination it's attached to, so a destination with the AWS Relations transformer writes only
aws_resource_relationships. To keep your raw AWS tables too, point the AWS Source at two PostgreSQL Destinations on the same database, one with the transformer and one without:kind: source
spec:
name: 'aws'
path: 'cloudquery/aws'
registry: 'cloudquery'
version: 'v34.9.0'
tables: ['aws_ec2_*', 'aws_iam_*', 'aws_s3_*', 'aws_account_information', 'aws_regions']
destinations: ['postgresql-raw', 'postgresql-relations']
spec: {}
---
kind: destination
spec:
name: 'postgresql-raw'
path: 'cloudquery/postgresql'
registry: 'cloudquery'
version: 'v8.16.0'
send_sync_summary: true
spec:
connection_string: '${PG_CONNECTION_STRING}'
---
kind: destination
spec:
name: 'postgresql-relations'
path: 'cloudquery/postgresql'
registry: 'cloudquery'
version: 'v8.16.0'
send_sync_summary: true
transformers: ['relations']
spec:
connection_string: '${PG_CONNECTION_STRING}'
---
kind: transformer
spec:
name: 'relations'
path: 'cloudquery/awsrelations'
registry: 'cloudquery'
version: 'v1.0.0'
spec: {}
Include
aws_account_information and aws_regions in the source's tables list, since every record gets an edge to its account and region. The optional skip_tables and skip_columns settings narrow what gets scanned; column matching is case-sensitive.What Doesn't Work Yet? #
This is a preview, and we'd rather you hear about the rough edges from us:
- Shared resources, like a shared VPC, don't join.
- Cross-account and cross-region ID references point at the wrong target.
- Versioned Lambda ARNs (
…:function:my-func:1) don't match the unqualified function row. - Tables keyed by two ARNs (
aws_ecs_cluster_services) or partly by a hash (aws_iam_policies) can't be targets. - RDS, DocumentDB and Neptune subnet groups all resolve to
aws_rds_subnet_groupsfor now.
The documentation has the details.
What Would You Build With This? #
We designed the transformer around questions we hear a lot: what breaks if I delete this, what depends on this role, what's attached to this VPC. If you run the CloudQuery CLI, try it on one of your accounts and email [email protected] with:
- The question you tried to answer: blast radius, incident investigation, access reviews, orphaned resource cleanup or something else.
- Whether the edges matched what you expected.
- What you'd want next: other clouds, a ready-made
aws_resourcesview, a graph export, or something else. - Which limitation causes you the most trouble.
A two-line email is plenty. The transformer is free during the preview and becomes a premium offering at general availability, so your feedback now shapes what ships then.
Frequently Asked Questions #
What table does the AWS Relations transformer create? #
A single table,
aws_resource_relationships, with fixed name and columns.Does the AWS Relations transformer keep my raw AWS tables? #
Not on the destination it's attached to. Use a second destination without the transformer, as shown above.
Which tables does the transformer scan? #
Every
aws_* table the AWS Source sends. Narrow the source's tables list or set skip_tables to scan less. Skipped tables can still be relationship targets.Why doesn't the transformer create edges from IAM policy documents? #
A policy says what a permission is scoped to, not what a resource is attached to, and a wildcard would connect a role to every bucket in the account. Roles still get edges to their attached and inline policies.
Do I have to sync the referenced table to get an edge? #
No. Edges are written even when the referenced table isn't synced. They won't join to anything, but
target_identifier still tells you what was referenced.Try CloudQuery yourself
Self-hosted, open source, and free to start. Or schedule a demo to see the managed version.