Complete Guide to GraphQL Federation
Build unified GraphQL APIs across multiple services with Apollo Federation. Covers subgraphs, supergraph composition, entity resolution, and gateway deployment.
Introduction
Building a single GraphQL schema for a whole company quickly becomes a bottleneck. Teams block each other on schema changes, deployments get coupled, and the monolithic graph becomes fragile. GraphQL Federation solves this by splitting the schema into subgraphs that each team owns and composing them back into one supergraph. This guide shows how to set up subgraphs, compose the supergraph, resolve entities, and deploy a gateway with Apollo Federation.
I’ve worked with GraphQL since 2018, and I’ve watched the monolithic pattern fail in three different companies. The story always goes the same way: one team needs to add a field, another team is mid-migration, and the schema review meeting turns into a two-hour argument. Federation doesn’t fix organizational problems, but it gives teams clear ownership boundaries. Each team owns its subgraph, deploys it on its own schedule, and the gateway composes them into a single API surface for clients.
If you’re coming from REST, the mental model is different. Instead of building one API per service and letting clients orchestrate calls, you compose one unified graph. Clients query one endpoint and the gateway figures out which subgraphs to call. For a deeper comparison, see GraphQL vs REST: A Complete Guide. If you’re already running microservices, federation fits naturally on top. Check out the Complete Guide to Microservices Communication for context on how services interact.
This guide assumes you know GraphQL basics: schemas, resolvers, and queries. You don’t need prior Apollo experience, but familiarity with Node.js and Python helps because the examples use both. By the end, you’ll have a working federated graph with three subgraphs, a gateway, and a supergraph composed with Rover. We’ll cover the architecture, subgraph setup in both Node.js and Python, gateway configuration, supergraph composition, entity resolution with all the key directives, querying patterns, best practices, and common mistakes. If you’re looking for production-specific guidance on deployment, monitoring, and managed federation, check out the Complete Guide to GraphQL Federation in Production.
Federation Architecture
A subgraph is a GraphQL service owned by a team that defines part of the schema. The supergraph is the composed schema built from all subgraphs. The gateway is the entry point that routes each part of a query to the right subgraph. An entity is a shared type with a key field that several subgraphs can reference and extend.
The key insight is that subgraphs don’t call each other directly. The gateway builds a query plan, calls the subgraphs in the right order, and joins the results. Subgraphs only need to know about entities they reference, not about the full graph. This keeps coupling low and lets teams work independently.
There are two composition models: managed federation (Apollo Studio hosts the supergraph schema and the gateway fetches it) and unmanaged federation (you compose the supergraph locally and the gateway loads it from a file or URL). Managed federation is better for production because it tracks schema changes, validates composition in CI, and lets you roll back. Unmanaged is fine for development and small deployments. I use managed federation in production and unmanaged for local development.
Subgraph Setup
Each subgraph is a standalone GraphQL service. It defines its own types,
queries, and mutations. The only federation-specific bits are the directives
(@key, @external, @extends, @requires, @provides, @shareable) that
tell the composition engine how types relate across subgraphs. The Apollo
Federation
specification
defines these directives. I remember being confused by them at first, but
they’re just annotations on your schema. You don’t change how resolvers work;
you add metadata that the gateway reads during composition.
Users subgraph (Node.js)
The Users subgraph owns the User type. It marks User as an entity with
@key(fields: "id"), which means other subgraphs can reference a user by its
id without needing to know how to resolve it. The subgraph also extends
Order and Product to add relationships from the user’s perspective. In
practice, this means the Users team controls what a User looks like, and
other teams can stitch their data onto it.
const { buildSubgraphSchema } = require("@apollo/subgraph");
const { gql, ApolloServer } = require("apollo-server");
const typeDefs = gql`
type User @key(fields: "id") {
id: ID!
name: String!
email: String!
orders: [Order!]!
}
extend type Order @key(fields: "id") {
id: ID! @external
user: User! @provides(fields: "name")
}
extend type Product @key(fields: "id") {
id: ID! @external
}
type Query {
user(id: ID!): User
users: [User!]!
}
`;
const resolvers = {
User: {
orders(user) {
return fetch(`http://orders-service/orders?userId=${user.id}`)
.then((res) => res.json());
},
},
Query: {
user: (_, { id }) => fetch(`http://users-service/users/${id}`).then((res) => res.json()),
users: () => fetch("http://users-service/users").then((res) => res.json()),
},
};
const server = new ApolloServer({
schema: buildSubgraphSchema([{ typeDefs, resolvers }]),
});
server.listen({ port: 4001 }).then(({ url }) => {
console.log(`Users subgraph ready at ${url}`);
});
Orders subgraph (Node.js)
The Orders subgraph owns the Order type and its fields. It references User
as an entity by returning a reference object { __typename: "User", id: order.userId } instead of fetching the full user. The gateway resolves the
user fields by calling the Users subgraph with that entity key. This is the
core of federation: subgraphs return entity references, and the gateway
follows them across subgraph boundaries.
The Order.user resolver returns a reference object, not a full user. This
is intentional. The gateway will call the Users subgraph to fill in the user
fields. If the client only asks for the order’s id and total, the gateway
never calls the Users subgraph at all. This lazy resolution is what makes
federation efficient.
const { buildSubgraphSchema } = require("@apollo/subgraph");
const { gql, ApolloServer } = require("apollo-server");
const typeDefs = gql`
type Order @key(fields: "id") {
id: ID!
total: Float!
status: String!
userId: ID!
user: User!
items: [OrderItem!]!
}
type OrderItem {
productId: ID!
quantity: Int!
price: Float!
}
extend type User @key(fields: "id") {
id: ID! @external
orders: [Order!]! @external
}
extend type Product @key(fields: "id") {
id: ID! @external
orders: [OrderItem!]!
}
type Query {
order(id: ID!): Order
orders: [Order!]!
}
type Mutation {
createOrder(userId: ID!, items: [OrderItemInput!]!): Order!
}
input OrderItemInput {
productId: ID!
quantity: Int!
}
`;
const resolvers = {
Order: {
user(order) {
return { __typename: "User", id: order.userId };
},
items(order) {
return order.items;
},
},
Product: {
orders(product) {
return fetch(`http://orders-service/orders/items?productId=${product.id}`)
.then((res) => res.json());
},
},
Query: {
order: (_, { id }) => fetch(`http://orders-service/orders/${id}`).then((res) => res.json()),
orders: () => fetch("http://orders-service/orders").then((res) => res.json()),
},
Mutation: {
createOrder: (_, { userId, items }) => {
return fetch("http://orders-service/orders", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ userId, items }),
}).then((res) => res.json());
},
},
};
const server = new ApolloServer({
schema: buildSubgraphSchema([{ typeDefs, resolvers }]),
});
server.listen({ port: 4002 }).then(({ url }) => {
console.log(`Orders subgraph ready at ${url}`);
});
Products subgraph (Python)
The Products subgraph uses Python with Ariadne,
which supports federation through make_federated_schema. The
__resolve_reference resolver is what makes the Product entity work across
subgraph boundaries. When the gateway needs to resolve a product reference
from the Orders subgraph, it calls this resolver with the entity key and the
Products subgraph returns the full product.
I chose Ariadne for this example because it’s the Python federation library I
know best. If you’re using Strawberry,
the approach is similar: you define a @key directive on the type and
implement a resolve_reference class method. The federation protocol is
language-agnostic, so any GraphQL server that uses it works.
from ariadne import QueryType, make_federated_schema, ObjectType
from ariadne.asgi import GraphQL
import httpx
type_defs = """
type Product @key(fields: "id") {
id: ID!
name: String!
price: Float!
description: String
}
type Query {
product(id: ID!): Product
products: [Product!]!
}
"""
query = QueryType()
product_obj = ObjectType("Product")
@query.field("product")
async def resolve_product(_, info, id):
async with httpx.AsyncClient() as client:
resp = await client.get(f"http://products-service/products/{id}")
return resp.json()
@query.field("products")
async def resolve_products(_, info):
async with httpx.AsyncClient() as client:
resp = await client.get("http://products-service/products")
return resp.json()
@product_obj.field("__resolve_reference")
async def resolve_product_reference(reference, info):
async with httpx.AsyncClient() as client:
resp = await client.get(f"http://products-service/products/{reference['id']}")
return resp.json()
schema = make_federated_schema(type_defs, [query, product_obj])
app = GraphQL(schema, debug=True)
Gateway Setup
The gateway is the single entry point for clients. It receives queries,
builds a query plan, calls the subgraphs, and joins the results. The
serviceList configuration tells the gateway where each subgraph lives. In
production, you’d use managed federation instead of a hardcoded service list
so the gateway can pick up schema changes without redeploying.
const { ApolloGateway } = require("@apollo/gateway");
const { ApolloServer } = require("apollo-server");
const gateway = new ApolloGateway({
serviceList: [
{ name: "users", url: "http://localhost:4001/graphql" },
{ name: "orders", url: "http://localhost:4002/graphql" },
{ name: "products", url: "http://localhost:4003/graphql" },
],
debug: true,
});
const server = new ApolloServer({
gateway,
subscriptions: false,
});
server.listen({ port: 4000 }).then(({ url }) => {
console.log(`Gateway ready at ${url}`);
});
The gateway handles query planning automatically. When a client asks for a
user and their orders, the gateway first calls the Users subgraph, extracts the
user’s id, then calls the Orders subgraph with that id as an entity key. It
joins the results and returns one response. The client never sees the
subgraph boundaries.
For production, consider the Apollo Router, a Rust-based gateway that’s faster than the Node.js Apollo Gateway. It supports the same federation protocol but handles higher throughput with lower latency. I’ve used both: the Node.js gateway is fine for development and small deployments, and the Router is better when you need to handle thousands of queries per second.
Supergraph Composition
Use the Rover CLI to compose the supergraph schema from the running subgraphs. Composition takes all subgraph schemas, validates them against the federation spec, and produces a single supergraph schema that the gateway uses to plan queries.
# Install Rover
brew install apollo-tooling/tap/rover
# Compose supergraph from subgraph schemas
rover supergraph compose --config supergraph.yaml > supergraph.graphql
# supergraph.yaml
federation_version: =2.8.0
subgraphs:
users:
routing_url: http://localhost:4001/graphql
schema:
subgraph_url: http://localhost:4001/graphql
orders:
routing_url: http://localhost:4002/graphql
schema:
subgraph_url: http://localhost:4002/graphql
products:
routing_url: http://localhost:4003/graphql
schema:
subgraph_url: http://localhost:4003/graphql
Composition can fail if two subgraphs define the same field without
@shareable, or if a subgraph uses @requires on a field that isn’t marked
@external. These errors surface in CI, which is where you want them. I run
rover supergraph compose in every CI pipeline so schema conflicts never reach
production. The Apollo documentation on
composition
covers the full list of composition rules and error messages.
In managed federation, you publish subgraph schemas to Apollo Studio with
rover subgraph publish instead of composing locally. Studio composes the
supergraph, validates it, and pushes it to the gateway. I’d recommend this
for any production setup because you get a schema history, composition
validation, and one-click rollback. The first time a composition breaks in
production at 2am, you’ll be glad you’ve got that rollback button.
Entity Resolution
Entities are the core of federation. They let one subgraph reference a type
owned by another subgraph without duplicating its definition. When the gateway
needs to resolve an entity reference, it calls the owning subgraph’s
__resolveReference resolver with the entity key. The subgraph returns the
full object, and the gateway fills in the fields the client requested.
I think of entities as the join tables of federation. In a monolithic schema,
you’d resolve a user’s orders with a single resolver that’s got access to both
the users and orders data sources. In federation, the Users subgraph returns
a User entity, the Orders subgraph extends it with orders, and the gateway
stitches them together. The subgraphs never talk to each other directly.
@key: define an entity
type User @key(fields: "id") {
id: ID!
name: String!
}
@extends: extend an entity from another subgraph
extend type User @key(fields: "id") {
id: ID! @external
orders: [Order!]!
}
@requires: compute fields based on external fields
extend type Product @key(fields: "id") {
id: ID! @external
price: Float! @external
discountedPrice: Float! @requires(fields: "price")
}
@provides: indicate a subgraph can provide fields of another type
extend type Order @key(fields: "id") {
id: ID! @external
user: User! @provides(fields: "name")
}
@shareable: mark a field as resolvable by several subgraphs
type Product @key(fields: "id") {
id: ID! @shareable
name: String! @shareable
}
@shareable tells the composition engine that two or more subgraphs can resolve
this field. Use it sparingly. If every subgraph can resolve every field, you’ve
lost the ownership boundary that federation is supposed to enforce. I use
@shareable only for fields that genuinely need to come from two or more sources,
like a Product.id that both the Products and Orders subgraphs need to return.
@inaccessible: hide a field from the public API
type User @key(fields: "id") {
id: ID!
internalNotes: String @inaccessible
}
@inaccessible lets you define a field in a subgraph without exposing it to
clients. The gateway strips it from the public schema. I use this for internal
fields that subgraphs need for @requires computations but shouldn’t be
queryable by clients.
Federation 2 vs Federation 1
Federation 2 simplified the directive model. In Federation 1, you needed
@extends on every extended type and @external on every foreign field.
Federation 2 made @extends optional and introduced @shareable and
@inaccessible. If you’re starting fresh, use Federation 2. If you’re on
Federation 1, the migration
guide
covers the changes. I migrated a federated graph from v1 to v2 and the main
benefit was less boilerplate in the subgraph schemas.
Querying the Federated Graph
This query spans all three subgraphs. The gateway sends the user portion to the
Users subgraph, uses the id to fetch orders from the Orders subgraph, and
resolves each product from the Products subgraph.
query GetUserWithOrders {
user(id: "1") {
id
name
email
orders {
id
total
status
items {
quantity
product {
name
price
}
}
}
}
}
The gateway builds a query plan that looks roughly like this: call Users to
get the user, extract the id, call Orders with that id to get the orders,
extract the productId from each order item, call Products with those IDs to
get the product names and prices. The client sees a single response, as if it
had queried a monolithic API.
If the Products subgraph is down, the gateway can return partial data: the
user and orders come through, but the product field returns null with an
error extension. This graceful degradation is one of federation’s biggest
wins over a monolith, where a single resolver failure can break the entire
query. For more on testing federated schemas, see the Complete Guide to
GraphQL Testing.
Best Practices
Keep one subgraph per team so ownership boundaries match organizational
boundaries. Use @key on any type that more than one subgraph needs to
reference. Keep each subgraph self-contained enough to run and test on its own.
I’ve seen teams try to split a single service into five subgraphs because they
thought more subgraphs meant more flexibility. It doesn’t. It means more
infrastructure, more composition errors, and more query plan complexity. Start
with two or three subgraphs and split only when a team genuinely needs
independent deployment.
Mark foreign fields with @external instead of redefining them. Avoid circular
extensions where two subgraphs keep referencing each other. Compose the
supergraph with Rover before deploying so schema conflicts surface in CI, not in
production. I run rover supergraph compose as a CI step on every pull request
that touches a subgraph schema. If composition fails, the PR stays blocked. This
catches issues like duplicate fields without @shareable before they reach
staging.
Cache entity resolution in the gateway, because __resolveReference runs
frequently. Monitor query plans to understand how a single client query turns
into several subgraph calls. If you use Apollo Studio, managed federation helps
track schema changes and composition errors across environments. The query
plan viewer in Studio is invaluable for spotting N+1 patterns early.
Version subgraphs independently; the gateway handles composition. When a subgraph fails, design the gateway to return partial data and error extensions instead of failing the whole request. Set timeouts on every subgraph call so one slow service doesn’t block the entire query. I learned this the hard way when a slow Products subgraph caused every query that touched products to time out, even though the user and order data was ready in 50ms. A 2-second timeout on subgraph calls fixed it.
When a field has to disappear from the supergraph, deprecate it in the owning subgraph and drive the removal through a documented timeline — the GraphQL Deprecation Policy Template is a copy-paste policy for exactly that, including usage tracking and removal criteria.
Use DataLoader for entity batching. Without it, resolving a list of 50 orders triggers 50 separate calls to the Products subgraph. DataLoader batches those into one call. This is the single biggest performance win in federation. If you do nothing else after reading this guide, add DataLoader to your entity resolvers.
Common Mistakes
Defining the same field in several subgraphs without @shareable will fail
composition. Forgetting to implement __resolveReference leaves entity lookups
returning null. Tight coupling between subgraphs defeats the purpose of
federation, because teams start depending on each other’s internals again. I
once reviewed a PR where a developer added a User.email field to the Orders
subgraph “just to avoid an extra call.” That’s exactly what federation is
designed to prevent. The Orders subgraph should return an entity reference and
let the Users subgraph resolve the email.
Not handling subgraph downtime means the gateway returns an error instead of
partial data. Using @requires on a field that isn’t marked @external fails
validation. Skipping local composition testing lets schema conflicts reach
production. Run rover supergraph compose locally before pushing to CI. It
takes 5 seconds and catches most composition errors.
Overusing @shareable blurs ownership boundaries. Ignoring query plan
performance can turn one query into an N+1 sequence of entity resolutions.
Exposing internal IDs across subgraph boundaries leaks implementation details.
Finally, not using DataLoader for entity batching can make a single client query
trigger hundreds of subgraph calls. I profiled a federated query once that made
127 subgraph calls for a list of 50 orders. Each order triggered a separate
Products subgraph call. Switching to DataLoader batched them into one call and
cut the query time from 3 seconds to 200ms. The fix took 15 minutes and saved
every query that touched products from that day on.
Summary
GraphQL Federation splits a monolithic schema into subgraphs that teams own
and deploy independently. The gateway composes them into a single API, builds
query plans, and routes each part of a query to the right subgraph. Entities
(@key, @extends, @external) are the glue that lets subgraphs reference
each other without coupling. Use Rover to compose the supergraph in CI, cache
entity resolution in the gateway, and set timeouts on every subgraph call. If
you remember nothing else: one subgraph per team, entities for cross-subgraph
references, and partial data over total failure.
I’ve been using federation in production for over three years now. The biggest benefit isn’t technical, it’s organizational. Teams can ship schema changes without coordinating with every other team. The biggest cost is operational: you’re now running three or more GraphQL services instead of one. Start small, with two or three subgraphs, and split more only when the team structure demands it. I started with two subgraphs and grew to seven over two years as the team grew.
See Also
- GraphQL vs REST: A Complete Guide: when to choose GraphQL over REST for your API
- Complete Guide to Microservices Communication: how microservices talk to each other, including GraphQL federation
- GraphQL Federated Entity Pattern: the entity pattern in detail with more examples
- Complete Guide to GraphQL Testing: how to test federated and non-federated GraphQL APIs
- GraphQL Deprecation Policy Template: a copy-paste policy for sunsetting fields and enum values across subgraphs
- Apollo Federation documentation: official docs covering the full federation specification
- Apollo Router documentation: the Rust-based gateway for high-throughput production deployments
Frequently Asked Questions
What is the difference between schema stitching and federation?
Schema stitching combines schemas by hand with custom resolvers. Federation
uses a standardized protocol (@key, @extends, and __resolveReference) so
subgraphs declare their relationships declaratively. For new projects, federation
is the better choice because it's more maintainable and has better tooling.
I migrated a project from stitching to federation back in 2020 and never looked
back. The biggest win wasn't having to write custom merge resolvers for every type.
How does the gateway handle a query that spans multiple subgraphs?
The gateway builds a query plan. For a query fetching a user and their orders,
it first calls the Users subgraph, then uses the user's id as an entity key to
call the Orders subgraph. It joins the results and returns one response to the
client. The query plan shows up in Apollo Studio, which helps you understand
the cost of each query.
Can I use federation without Apollo?
Yes. Federation is an open spec. You can use Apollo Gateway (Node.js), Apollo Router (Rust), or a custom gateway. The subgraph protocol is language-agnostic, so subgraphs can be built in Python (Ariadne, Strawberry), Java (DGS), Go (gqlgen), and Ruby (graphql-ruby). I've mixed Node.js and Python subgraphs in the same federated graph without issues.
When should I prefer a monolithic GraphQL API over federation?
Federation pays off when several teams own different parts of the schema and need to deploy independently. If your API is small, has one owner, and few coupling points, a monolithic schema is simpler and has less overhead. I usually recommend federation when you've got three or more teams contributing to the same GraphQL API. Below that, the infrastructure overhead isn't worth it.
How do I handle authentication in a federated graph?
Handle auth at the gateway level, not in each subgraph. The gateway validates the token, extracts the user context, and passes it to subgraphs via request headers. Subgraphs trust the gateway and use the context to authorize access. This avoids duplicating auth logic across subgraphs and keeps the gateway as the single enforcement point.
Related Resources
GraphQL vs REST — When to Choose and How to Migrate
A decision guide comparing GraphQL and REST APIs: use cases, performance, caching, tooling, and migration strategies for engineering teams.
GuideComplete Guide to API Versioning Strategies
Version REST and GraphQL APIs with URI, header, query param, and content negotiation strategies. Covers deprecation, sunset, and migration patterns.
GuideComplete Guide to Microservices Communication
Compare sync vs async communication patterns for microservices. Covers REST, gRPC, message queues, event-driven, service mesh, and when to use each.
PatternGraphQL Federated Entity Pattern
Share an entity across Apollo Federation subgraphs. Use @key, @external, and @shareable so each service owns the fields it knows best.
GuideGraphQL Federation in Production
Run federated GraphQL in production with confidence. Covers subgraph composition, gateway deployment, entity resolution, schema coordination, observability, and failure handling.
GuideComplete Guide to GraphQL Testing
Test GraphQL APIs at every layer: unit tests for resolvers, integration tests for schema, E2E tests for operations. Covers mocking, fixtures, snapshot testing, and performance testing.