StackPractices
advanced By Mathias Paulenko

Complete Guide to GraphQL Federation

Build unified GraphQL APIs across multiple services with Apollo Federation. Covers subgraphs, supergraph composition, entity resolution, and gateway deployment.

Introduction

Building a single GraphQL schema for a whole company quickly becomes a bottleneck. Teams block each other on schema changes, deployments get coupled, and the monolithic graph becomes fragile. GraphQL Federation solves this by splitting the schema into subgraphs that each team owns and composing them back into one supergraph. This guide shows how to set up subgraphs, compose the supergraph, resolve entities, and deploy a gateway with Apollo Federation.

I’ve worked with GraphQL since 2018, and I’ve watched the monolithic pattern fail in three different companies. The story always goes the same way: one team needs to add a field, another team is mid-migration, and the schema review meeting turns into a two-hour argument. Federation doesn’t fix organizational problems, but it gives teams clear ownership boundaries. Each team owns its subgraph, deploys it on its own schedule, and the gateway composes them into a single API surface for clients.

If you’re coming from REST, the mental model is different. Instead of building one API per service and letting clients orchestrate calls, you compose one unified graph. Clients query one endpoint and the gateway figures out which subgraphs to call. For a deeper comparison, see GraphQL vs REST: A Complete Guide. If you’re already running microservices, federation fits naturally on top. Check out the Complete Guide to Microservices Communication for context on how services interact.

This guide assumes you know GraphQL basics: schemas, resolvers, and queries. You don’t need prior Apollo experience, but familiarity with Node.js and Python helps because the examples use both. By the end, you’ll have a working federated graph with three subgraphs, a gateway, and a supergraph composed with Rover. We’ll cover the architecture, subgraph setup in both Node.js and Python, gateway configuration, supergraph composition, entity resolution with all the key directives, querying patterns, best practices, and common mistakes. If you’re looking for production-specific guidance on deployment, monitoring, and managed federation, check out the Complete Guide to GraphQL Federation in Production.

Federation Architecture

flowchart diagram: Client

A subgraph is a GraphQL service owned by a team that defines part of the schema. The supergraph is the composed schema built from all subgraphs. The gateway is the entry point that routes each part of a query to the right subgraph. An entity is a shared type with a key field that several subgraphs can reference and extend.

The key insight is that subgraphs don’t call each other directly. The gateway builds a query plan, calls the subgraphs in the right order, and joins the results. Subgraphs only need to know about entities they reference, not about the full graph. This keeps coupling low and lets teams work independently.

There are two composition models: managed federation (Apollo Studio hosts the supergraph schema and the gateway fetches it) and unmanaged federation (you compose the supergraph locally and the gateway loads it from a file or URL). Managed federation is better for production because it tracks schema changes, validates composition in CI, and lets you roll back. Unmanaged is fine for development and small deployments. I use managed federation in production and unmanaged for local development.

Subgraph Setup

Each subgraph is a standalone GraphQL service. It defines its own types, queries, and mutations. The only federation-specific bits are the directives (@key, @external, @extends, @requires, @provides, @shareable) that tell the composition engine how types relate across subgraphs. The Apollo Federation specification defines these directives. I remember being confused by them at first, but they’re just annotations on your schema. You don’t change how resolvers work; you add metadata that the gateway reads during composition.

Users subgraph (Node.js)

The Users subgraph owns the User type. It marks User as an entity with @key(fields: "id"), which means other subgraphs can reference a user by its id without needing to know how to resolve it. The subgraph also extends Order and Product to add relationships from the user’s perspective. In practice, this means the Users team controls what a User looks like, and other teams can stitch their data onto it.

const { buildSubgraphSchema } = require("@apollo/subgraph");
const { gql, ApolloServer } = require("apollo-server");

const typeDefs = gql`
  type User @key(fields: "id") {
    id: ID!
    name: String!
    email: String!
    orders: [Order!]!
  }

  extend type Order @key(fields: "id") {
    id: ID! @external
    user: User! @provides(fields: "name")
  }

  extend type Product @key(fields: "id") {
    id: ID! @external
  }

  type Query {
    user(id: ID!): User
    users: [User!]!
  }
`;

const resolvers = {
  User: {
    orders(user) {
      return fetch(`http://orders-service/orders?userId=${user.id}`)
        .then((res) => res.json());
    },
  },
  Query: {
    user: (_, { id }) => fetch(`http://users-service/users/${id}`).then((res) => res.json()),
    users: () => fetch("http://users-service/users").then((res) => res.json()),
  },
};

const server = new ApolloServer({
  schema: buildSubgraphSchema([{ typeDefs, resolvers }]),
});

server.listen({ port: 4001 }).then(({ url }) => {
  console.log(`Users subgraph ready at ${url}`);
});

Orders subgraph (Node.js)

The Orders subgraph owns the Order type and its fields. It references User as an entity by returning a reference object { __typename: "User", id: order.userId } instead of fetching the full user. The gateway resolves the user fields by calling the Users subgraph with that entity key. This is the core of federation: subgraphs return entity references, and the gateway follows them across subgraph boundaries.

The Order.user resolver returns a reference object, not a full user. This is intentional. The gateway will call the Users subgraph to fill in the user fields. If the client only asks for the order’s id and total, the gateway never calls the Users subgraph at all. This lazy resolution is what makes federation efficient.

const { buildSubgraphSchema } = require("@apollo/subgraph");
const { gql, ApolloServer } = require("apollo-server");

const typeDefs = gql`
  type Order @key(fields: "id") {
    id: ID!
    total: Float!
    status: String!
    userId: ID!
    user: User!
    items: [OrderItem!]!
  }

  type OrderItem {
    productId: ID!
    quantity: Int!
    price: Float!
  }

  extend type User @key(fields: "id") {
    id: ID! @external
    orders: [Order!]! @external
  }

  extend type Product @key(fields: "id") {
    id: ID! @external
    orders: [OrderItem!]!
  }

  type Query {
    order(id: ID!): Order
    orders: [Order!]!
  }

  type Mutation {
    createOrder(userId: ID!, items: [OrderItemInput!]!): Order!
  }

  input OrderItemInput {
    productId: ID!
    quantity: Int!
  }
`;

const resolvers = {
  Order: {
    user(order) {
      return { __typename: "User", id: order.userId };
    },
    items(order) {
      return order.items;
    },
  },
  Product: {
    orders(product) {
      return fetch(`http://orders-service/orders/items?productId=${product.id}`)
        .then((res) => res.json());
    },
  },
  Query: {
    order: (_, { id }) => fetch(`http://orders-service/orders/${id}`).then((res) => res.json()),
    orders: () => fetch("http://orders-service/orders").then((res) => res.json()),
  },
  Mutation: {
    createOrder: (_, { userId, items }) => {
      return fetch("http://orders-service/orders", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ userId, items }),
      }).then((res) => res.json());
    },
  },
};

const server = new ApolloServer({
  schema: buildSubgraphSchema([{ typeDefs, resolvers }]),
});

server.listen({ port: 4002 }).then(({ url }) => {
  console.log(`Orders subgraph ready at ${url}`);
});

Products subgraph (Python)

The Products subgraph uses Python with Ariadne, which supports federation through make_federated_schema. The __resolve_reference resolver is what makes the Product entity work across subgraph boundaries. When the gateway needs to resolve a product reference from the Orders subgraph, it calls this resolver with the entity key and the Products subgraph returns the full product.

I chose Ariadne for this example because it’s the Python federation library I know best. If you’re using Strawberry, the approach is similar: you define a @key directive on the type and implement a resolve_reference class method. The federation protocol is language-agnostic, so any GraphQL server that uses it works.

from ariadne import QueryType, make_federated_schema, ObjectType
from ariadne.asgi import GraphQL
import httpx

type_defs = """
    type Product @key(fields: "id") {
        id: ID!
        name: String!
        price: Float!
        description: String
    }

    type Query {
        product(id: ID!): Product
        products: [Product!]!
    }
"""

query = QueryType()
product_obj = ObjectType("Product")

@query.field("product")
async def resolve_product(_, info, id):
    async with httpx.AsyncClient() as client:
        resp = await client.get(f"http://products-service/products/{id}")
        return resp.json()

@query.field("products")
async def resolve_products(_, info):
    async with httpx.AsyncClient() as client:
        resp = await client.get("http://products-service/products")
        return resp.json()

@product_obj.field("__resolve_reference")
async def resolve_product_reference(reference, info):
    async with httpx.AsyncClient() as client:
        resp = await client.get(f"http://products-service/products/{reference['id']}")
        return resp.json()

schema = make_federated_schema(type_defs, [query, product_obj])
app = GraphQL(schema, debug=True)

Gateway Setup

The gateway is the single entry point for clients. It receives queries, builds a query plan, calls the subgraphs, and joins the results. The serviceList configuration tells the gateway where each subgraph lives. In production, you’d use managed federation instead of a hardcoded service list so the gateway can pick up schema changes without redeploying.

const { ApolloGateway } = require("@apollo/gateway");
const { ApolloServer } = require("apollo-server");

const gateway = new ApolloGateway({
  serviceList: [
    { name: "users", url: "http://localhost:4001/graphql" },
    { name: "orders", url: "http://localhost:4002/graphql" },
    { name: "products", url: "http://localhost:4003/graphql" },
  ],
  debug: true,
});

const server = new ApolloServer({
  gateway,
  subscriptions: false,
});

server.listen({ port: 4000 }).then(({ url }) => {
  console.log(`Gateway ready at ${url}`);
});

The gateway handles query planning automatically. When a client asks for a user and their orders, the gateway first calls the Users subgraph, extracts the user’s id, then calls the Orders subgraph with that id as an entity key. It joins the results and returns one response. The client never sees the subgraph boundaries.

For production, consider the Apollo Router, a Rust-based gateway that’s faster than the Node.js Apollo Gateway. It supports the same federation protocol but handles higher throughput with lower latency. I’ve used both: the Node.js gateway is fine for development and small deployments, and the Router is better when you need to handle thousands of queries per second.

Supergraph Composition

Use the Rover CLI to compose the supergraph schema from the running subgraphs. Composition takes all subgraph schemas, validates them against the federation spec, and produces a single supergraph schema that the gateway uses to plan queries.

# Install Rover
brew install apollo-tooling/tap/rover

# Compose supergraph from subgraph schemas
rover supergraph compose --config supergraph.yaml > supergraph.graphql
# supergraph.yaml
federation_version: =2.8.0
subgraphs:
  users:
    routing_url: http://localhost:4001/graphql
    schema:
      subgraph_url: http://localhost:4001/graphql
  orders:
    routing_url: http://localhost:4002/graphql
    schema:
      subgraph_url: http://localhost:4002/graphql
  products:
    routing_url: http://localhost:4003/graphql
    schema:
      subgraph_url: http://localhost:4003/graphql

Composition can fail if two subgraphs define the same field without @shareable, or if a subgraph uses @requires on a field that isn’t marked @external. These errors surface in CI, which is where you want them. I run rover supergraph compose in every CI pipeline so schema conflicts never reach production. The Apollo documentation on composition covers the full list of composition rules and error messages.

In managed federation, you publish subgraph schemas to Apollo Studio with rover subgraph publish instead of composing locally. Studio composes the supergraph, validates it, and pushes it to the gateway. I’d recommend this for any production setup because you get a schema history, composition validation, and one-click rollback. The first time a composition breaks in production at 2am, you’ll be glad you’ve got that rollback button.

Entity Resolution

Entities are the core of federation. They let one subgraph reference a type owned by another subgraph without duplicating its definition. When the gateway needs to resolve an entity reference, it calls the owning subgraph’s __resolveReference resolver with the entity key. The subgraph returns the full object, and the gateway fills in the fields the client requested.

I think of entities as the join tables of federation. In a monolithic schema, you’d resolve a user’s orders with a single resolver that’s got access to both the users and orders data sources. In federation, the Users subgraph returns a User entity, the Orders subgraph extends it with orders, and the gateway stitches them together. The subgraphs never talk to each other directly.

@key: define an entity

type User @key(fields: "id") {
  id: ID!
  name: String!
}

@extends: extend an entity from another subgraph

extend type User @key(fields: "id") {
  id: ID! @external
  orders: [Order!]!
}

@requires: compute fields based on external fields

extend type Product @key(fields: "id") {
  id: ID! @external
  price: Float! @external
  discountedPrice: Float! @requires(fields: "price")
}

@provides: indicate a subgraph can provide fields of another type

extend type Order @key(fields: "id") {
  id: ID! @external
  user: User! @provides(fields: "name")
}

@shareable: mark a field as resolvable by several subgraphs

type Product @key(fields: "id") {
  id: ID! @shareable
  name: String! @shareable
}

@shareable tells the composition engine that two or more subgraphs can resolve this field. Use it sparingly. If every subgraph can resolve every field, you’ve lost the ownership boundary that federation is supposed to enforce. I use @shareable only for fields that genuinely need to come from two or more sources, like a Product.id that both the Products and Orders subgraphs need to return.

@inaccessible: hide a field from the public API

type User @key(fields: "id") {
  id: ID!
  internalNotes: String @inaccessible
}

@inaccessible lets you define a field in a subgraph without exposing it to clients. The gateway strips it from the public schema. I use this for internal fields that subgraphs need for @requires computations but shouldn’t be queryable by clients.

Federation 2 vs Federation 1

Federation 2 simplified the directive model. In Federation 1, you needed @extends on every extended type and @external on every foreign field. Federation 2 made @extends optional and introduced @shareable and @inaccessible. If you’re starting fresh, use Federation 2. If you’re on Federation 1, the migration guide covers the changes. I migrated a federated graph from v1 to v2 and the main benefit was less boilerplate in the subgraph schemas.

Querying the Federated Graph

This query spans all three subgraphs. The gateway sends the user portion to the Users subgraph, uses the id to fetch orders from the Orders subgraph, and resolves each product from the Products subgraph.

query GetUserWithOrders {
  user(id: "1") {
    id
    name
    email
    orders {
      id
      total
      status
      items {
        quantity
        product {
          name
          price
        }
      }
    }
  }
}

The gateway builds a query plan that looks roughly like this: call Users to get the user, extract the id, call Orders with that id to get the orders, extract the productId from each order item, call Products with those IDs to get the product names and prices. The client sees a single response, as if it had queried a monolithic API.

If the Products subgraph is down, the gateway can return partial data: the user and orders come through, but the product field returns null with an error extension. This graceful degradation is one of federation’s biggest wins over a monolith, where a single resolver failure can break the entire query. For more on testing federated schemas, see the Complete Guide to GraphQL Testing.

Best Practices

Keep one subgraph per team so ownership boundaries match organizational boundaries. Use @key on any type that more than one subgraph needs to reference. Keep each subgraph self-contained enough to run and test on its own. I’ve seen teams try to split a single service into five subgraphs because they thought more subgraphs meant more flexibility. It doesn’t. It means more infrastructure, more composition errors, and more query plan complexity. Start with two or three subgraphs and split only when a team genuinely needs independent deployment.

Mark foreign fields with @external instead of redefining them. Avoid circular extensions where two subgraphs keep referencing each other. Compose the supergraph with Rover before deploying so schema conflicts surface in CI, not in production. I run rover supergraph compose as a CI step on every pull request that touches a subgraph schema. If composition fails, the PR stays blocked. This catches issues like duplicate fields without @shareable before they reach staging.

Cache entity resolution in the gateway, because __resolveReference runs frequently. Monitor query plans to understand how a single client query turns into several subgraph calls. If you use Apollo Studio, managed federation helps track schema changes and composition errors across environments. The query plan viewer in Studio is invaluable for spotting N+1 patterns early.

Version subgraphs independently; the gateway handles composition. When a subgraph fails, design the gateway to return partial data and error extensions instead of failing the whole request. Set timeouts on every subgraph call so one slow service doesn’t block the entire query. I learned this the hard way when a slow Products subgraph caused every query that touched products to time out, even though the user and order data was ready in 50ms. A 2-second timeout on subgraph calls fixed it.

When a field has to disappear from the supergraph, deprecate it in the owning subgraph and drive the removal through a documented timeline — the GraphQL Deprecation Policy Template is a copy-paste policy for exactly that, including usage tracking and removal criteria.

Use DataLoader for entity batching. Without it, resolving a list of 50 orders triggers 50 separate calls to the Products subgraph. DataLoader batches those into one call. This is the single biggest performance win in federation. If you do nothing else after reading this guide, add DataLoader to your entity resolvers.

Common Mistakes

Defining the same field in several subgraphs without @shareable will fail composition. Forgetting to implement __resolveReference leaves entity lookups returning null. Tight coupling between subgraphs defeats the purpose of federation, because teams start depending on each other’s internals again. I once reviewed a PR where a developer added a User.email field to the Orders subgraph “just to avoid an extra call.” That’s exactly what federation is designed to prevent. The Orders subgraph should return an entity reference and let the Users subgraph resolve the email.

Not handling subgraph downtime means the gateway returns an error instead of partial data. Using @requires on a field that isn’t marked @external fails validation. Skipping local composition testing lets schema conflicts reach production. Run rover supergraph compose locally before pushing to CI. It takes 5 seconds and catches most composition errors.

Overusing @shareable blurs ownership boundaries. Ignoring query plan performance can turn one query into an N+1 sequence of entity resolutions. Exposing internal IDs across subgraph boundaries leaks implementation details. Finally, not using DataLoader for entity batching can make a single client query trigger hundreds of subgraph calls. I profiled a federated query once that made 127 subgraph calls for a list of 50 orders. Each order triggered a separate Products subgraph call. Switching to DataLoader batched them into one call and cut the query time from 3 seconds to 200ms. The fix took 15 minutes and saved every query that touched products from that day on.

Summary

GraphQL Federation splits a monolithic schema into subgraphs that teams own and deploy independently. The gateway composes them into a single API, builds query plans, and routes each part of a query to the right subgraph. Entities (@key, @extends, @external) are the glue that lets subgraphs reference each other without coupling. Use Rover to compose the supergraph in CI, cache entity resolution in the gateway, and set timeouts on every subgraph call. If you remember nothing else: one subgraph per team, entities for cross-subgraph references, and partial data over total failure.

I’ve been using federation in production for over three years now. The biggest benefit isn’t technical, it’s organizational. Teams can ship schema changes without coordinating with every other team. The biggest cost is operational: you’re now running three or more GraphQL services instead of one. Start small, with two or three subgraphs, and split more only when the team structure demands it. I started with two subgraphs and grew to seven over two years as the team grew.

See Also

Frequently Asked Questions

What is the difference between schema stitching and federation?

Schema stitching combines schemas by hand with custom resolvers. Federation uses a standardized protocol (@key, @extends, and __resolveReference) so subgraphs declare their relationships declaratively. For new projects, federation is the better choice because it's more maintainable and has better tooling. I migrated a project from stitching to federation back in 2020 and never looked back. The biggest win wasn't having to write custom merge resolvers for every type.

How does the gateway handle a query that spans multiple subgraphs?

The gateway builds a query plan. For a query fetching a user and their orders, it first calls the Users subgraph, then uses the user's id as an entity key to call the Orders subgraph. It joins the results and returns one response to the client. The query plan shows up in Apollo Studio, which helps you understand the cost of each query.

Can I use federation without Apollo?

Yes. Federation is an open spec. You can use Apollo Gateway (Node.js), Apollo Router (Rust), or a custom gateway. The subgraph protocol is language-agnostic, so subgraphs can be built in Python (Ariadne, Strawberry), Java (DGS), Go (gqlgen), and Ruby (graphql-ruby). I've mixed Node.js and Python subgraphs in the same federated graph without issues.

When should I prefer a monolithic GraphQL API over federation?

Federation pays off when several teams own different parts of the schema and need to deploy independently. If your API is small, has one owner, and few coupling points, a monolithic schema is simpler and has less overhead. I usually recommend federation when you've got three or more teams contributing to the same GraphQL API. Below that, the infrastructure overhead isn't worth it.

How do I handle authentication in a federated graph?

Handle auth at the gateway level, not in each subgraph. The gateway validates the token, extracts the user context, and passes it to subgraphs via request headers. Subgraphs trust the gateway and use the context to authorize access. This avoids duplicating auth logic across subgraphs and keeps the gateway as the single enforcement point.