Skip to main content

Command Palette

Search for a command to run...

Your Test Suite Isn't Ready for AI Agents — Here's How to Fix That

Updated
•8 min read•View as Markdown
Your Test Suite Isn't Ready for AI Agents — Here's How to Fix That
N
Love to code, gaming. And I use vim btw.

Momentic just shipped Mo — an AI agent that generates automated tests without you writing a single line of test scaffolding. It landed yesterday, and it's one of a growing wave of tools that want to take QA off your plate entirely. That's genuinely exciting. But here's the thing nobody's saying out loud: most MERN codebases are so poorly structured that handing them to an AI testing agent would just produce 200 flaky tests you'd spend the rest of the sprint babysitting. The tool isn't the problem. Your architecture is.

Let's fix that.

The Testing Pyramid Is Still the Law

I know this feels like being told to eat your vegetables. But the reason we keep coming back to it is because most teams still skip it, and then wonder why their CI pipeline takes 45 minutes and still misses bugs.

The testing pyramid breaks down like this:

  • Unit tests at the bottom — a lot of them, fast, cheap, isolated from everything. These should run in under a second each.

  • Integration tests in the middle — fewer, testing how your pieces actually talk to each other (think Express + MongoDB).

  • E2E tests at the top — the smallest layer. Real user flows through a browser. Slow and expensive, use sparingly.

For a MERN stack this translates pretty directly. Unit tests live at the service/controller layer on the Node side, and at the component/hook level on the React side. Integration tests live where your Express routes hit your MongoDB models. E2E tests cover your two or three most important user journeys — login, checkout, whatever makes you money.

Most MERN devs do the inverse of this. They write one Cypress test that covers the whole signup flow, call it "tested," and ship. Then they wonder why every refactor breaks something in a way that's painful to debug.

Unit Testing Express Routes Without Losing Your Mind

The reason people hate unit testing their Express handlers isn't because unit testing is hard — it's because they wrote routes that are impossible to isolate. Look at this:

// ❌ You can't test this without a live database
app.get('/api/users/:id', async (req, res) => {
  try {
    const user = await User.findById(req.params.id);
    if (!user) return res.status(404).json({ message: 'Not found' });
    res.json(user);
  } catch (err) {
    res.status(500).json({ message: 'Server error' });
  }
});

The database call is baked directly into the route. There's no seam to inject a mock. Every test you write for this needs MongoDB running, which makes it slow, fragile, and annoying to set up in CI.

Here's the version that's actually testable:

// ✅ controllers/userController.js — testable by design
export const getUserById = async (req, res, { userService }) => {
  try {
    const user = await userService.findById(req.params.id);
    if (!user) return res.status(404).json({ message: 'Not found' });
    res.json(user);
  } catch (err) {
    res.status(500).json({ message: 'Server error' });
  }
};
// user.test.js — pure Jest, zero infrastructure
import { getUserById } from '../controllers/userController';

describe('getUserById', () => {
  it('returns 404 when user does not exist', async () => {
    const req = { params: { id: 'abc123' } };
    const res = { status: jest.fn().mockReturnThis(), json: jest.fn() };
    const userService = { findById: jest.fn().mockResolvedValue(null) };

    await getUserById(req, res, { userService });

    expect(res.status).toHaveBeenCalledWith(404);
    expect(res.json).toHaveBeenCalledWith({ message: 'Not found' });
  });

  it('returns user data on success', async () => {
    const mockUser = { _id: 'abc123', name: 'Yuu' };
    const req = { params: { id: 'abc123' } };
    const res = { status: jest.fn().mockReturnThis(), json: jest.fn() };
    const userService = { findById: jest.fn().mockResolvedValue(mockUser) };

    await getUserById(req, res, { userService });

    expect(res.json).toHaveBeenCalledWith(mockUser);
  });
});

This runs in milliseconds. No database, no Docker, no headaches. And when Mo or any AI agent looks at this code, it has a clear surface to work with.

Integration Tests: mongodb-memory-server Is Non-Negotiable

Unit tests are great, but they won't tell you if your Mongoose schema validation is wired up right, or if your indexes are actually enforcing uniqueness. That's what integration tests are for — and the right way to run them is with an in-memory MongoDB instance.

npm install --save-dev mongodb-memory-server
// tests/integration/user.integration.test.js
import { MongoMemoryServer } from 'mongodb-memory-server';
import mongoose from 'mongoose';
import request from 'supertest';
import app from '../../app';
import { User } from '../../models/User';

let mongod;

beforeAll(async () => {
  mongod = await MongoMemoryServer.create();
  await mongoose.connect(mongod.getUri());
});

afterAll(async () => {
  await mongoose.disconnect();
  await mongod.stop();
});

afterEach(async () => {
  await User.deleteMany({});
});

describe('POST /api/users', () => {
  it('creates a user and returns 201', async () => {
    const response = await request(app)
      .post('/api/users')
      .send({ name: 'Yuu', email: '[email protected]', password: 'secure123' });

    expect(response.status).toBe(201);
    expect(response.body.email).toBe('[email protected]');
    // password should be hashed, not stored plaintext
    expect(response.body.password).toBeUndefined();
  });

  it('rejects duplicate emails with 409', async () => {
    await User.create({ name: 'Yuu', email: '[email protected]', password: 'hashed' });

    const response = await request(app)
      .post('/api/users')
      .send({ name: 'Yuu', email: '[email protected]', password: 'secure123' });

    expect(response.status).toBe(409);
  });
});

This is the real deal — it hits your actual Mongoose model, runs your validators, tests your middleware chain. It's slower than a unit test but still much faster than a full E2E, and it catches a completely different class of bugs.

React: Test What the User Does, Not What Your Code Does

React Testing Library made one thing very clear when it launched: stop testing state. Stop testing internal methods. Stop testing whether useState was called. Test what a user would actually experience.

Here's a before and after for a login form:

// ❌ Testing implementation — brittle and pointless
test('sets loading state on submit', () => {
  const { getByRole } = render(<LoginForm />);
  fireEvent.click(getByRole('button'));
  // asserting on internal loading state directly
  // this breaks every time you rename a state variable
});
// ✅ Testing user behavior — this is what actually matters
import { render, screen, fireEvent, waitFor } from '@testing-library/react';
import { rest } from 'msw';
import { setupServer } from 'msw/node';
import LoginForm from '../components/LoginForm';

const server = setupServer(
  rest.post('/api/auth/login', (req, res, ctx) => {
    return res(ctx.status(401), ctx.json({ message: 'Invalid credentials' }));
  })
);

beforeAll(() => server.listen());
afterEach(() => server.resetHandlers());
afterAll(() => server.close());

test('shows error message when credentials are wrong', async () => {
  render(<LoginForm />);

  fireEvent.change(screen.getByLabelText(/email/i), {
    target: { value: '[email protected]' },
  });
  fireEvent.change(screen.getByLabelText(/password/i), {
    target: { value: 'wrongpassword' },
  });
  fireEvent.click(screen.getByRole('button', { name: /sign in/i }));

  await screen.findByText(/invalid credentials/i);
});

The key here is msw (Mock Service Worker) — it intercepts the actual HTTP calls your component makes and returns controlled responses. Your test doesn't care about axios, fetch, or whatever you're using internally. It just cares about what shows up on screen.

Before You Hand Mo the Keys

If you want an AI testing agent to produce tests that are actually useful, here's the quick checklist to run through first:

Architecture: Is your logic separated from your I/O? Can you call a function without needing a database, an HTTP server, or a filesystem? If not, start there — no AI can save a tightly coupled codebase.

Naming: Are your files and test descriptions readable? describe('getUserById when user does not exist') tells Mo (and future you) exactly what's being tested. test('test 2') tells nobody anything.

Isolation: Do your tests share state? A test that relies on another test having run first will silently fail in ways that take hours to debug. Every test should set up what it needs and tear down what it created.

Coverage of what matters: You don't need 100% coverage. You need solid coverage of the paths that can actually break in production — validation logic, auth flows, error handling, edge cases around your data models.

Get these four things right and Mo becomes a genuine multiplier. Hand it a mess and you'll get more mess, just generated faster.

Where This Is All Heading

The next 12 months are going to see a lot more tools like Mo. The boring, repeatable parts of software development — boilerplate tests, documentation, code review — are exactly the kind of tasks AI agents are good at. That's not a threat to your job. It's a shift in what your job looks like.

The developers who benefit from AI testing agents aren't the ones who hand over a spaghetti codebase and hope for the best. They're the ones who've already built clean, testable code and now get to skip the grunt work of writing test scaffolding.

Write the architecture first. Let the AI fill in the rest.