In your specs
def call
gateway.charge(cart.total)
end allow(gateway)
.to receive(:charge) def charge(amount)
…
end In production
The problem
Coverage counts lines, not collaborations. A suite can cover every line of two classes and still never run them together.
Checkout’s spec stubs the payment gateway, because it should: the unit spec is about Checkout.
PaymentGateway’s spec tests the gateway on its own. Both files are at 100% coverage, and no example
anywhere runs the real Checkout calling the real PaymentGateway#charge. If the two disagree about
what #charge takes or returns, every test stays green.
This happens whether you write tests first or after, and it happens most in a disciplined, mock-heavy suite. Every unit is isolated, which is the point, and nothing checks that the isolated units fit.
The fix
$ ra testing_pyramid --stubbed-only
212 boundaries · only crossed via doubles
Checkout#call → PaymentGateway#charge
3 stubbed · 0 real · 8,204 prod
Report#build → Invoice#total
2 stubbed · 0 real · 4,102 prod
Each boundary shows how many specs stub it, how many cross it for real, and, when you’ve recorded production, how often it runs there. Without production data the list is ordered by how many specs depend on the stub.
Then it writes the test that crosses the seam, with setup taken from a real run:
# setup taken from a real run (how_to_reach)
RSpec.describe Checkout do
it "charges the gateway" do
gateway = PaymentGateway.new(account:, rate_table:)
checkout = Checkout.new(cart:, gateway:)
expect { checkout.call }.to change(gateway, :charged?).to(true)
end
end
How it works
- Collect the boundaries. Every recorded call from one class’s method into another’s is an edge.
Every stub in a spec is one too: a spec whose subject is
Checkoutstubbing#chargeon aPaymentGatewaydouble names the edgeCheckout#call → PaymentGateway#charge. - Label each crossing. Each call tree is stamped with the spec example that produced it, so every real crossing knows which spec made it, and which layer that spec belongs to (unit, integration or end-to-end, from its type or directory).
- Find the gaps. An edge with stubs but no real crossing, at any layer, has only ever been crossed through a double.
- Weight them. With production or staging recorded, each gap gets its real call count, so the busiest untested seams come first.
Limits
- It has to know the edge exists. A seam shows up if a spec stubs it or if a recorded run crossed it. A collaboration that nothing stubs and nothing ran is invisible.
- Layers come from your conventions. It reads layer from spec type metadata or directory
(
spec/models,spec/requests,spec/system). An unconventional layout needs mapping once. - The scaffold is a start. The setup comes from a real run. The assertion that matters is still yours to write.
Related tools
verify_mock
Catch a lying mock. A stub that passes green but returns something the real collaborator never would. Mutation testing can't see it. The recording can.
See more →how_to_reach
Reach any method. The collaborators and call path that get an object into the state a method needs, taken from a real run. Setup for tests that use real objects instead of doubles.
See more →specs_that_reach
Run only the specs that matter. The spec examples whose recorded runs reach a method, so a change reruns those and nothing else. Keeps a red-green loop in seconds.
See more →