At Easygo, our products run across a growing number of markets, and we're rapidly expanding into more. Back when it was just two markets though, our Playwright E2E suite had one folder per market. E2E tests, or end-to-end, click through the site the way an actual person would and catch things before an actual customer runs into them.
This approach made sense at the time, or so we thought. Except the same functionality living in two markets just meant the same test living twice, one copy sitting in each market's folder, copy and pasted, and then slowly drifting apart from each other over time.
So we rebuilt the test framework around Critical user journeys instead, the ones defined by our product squads. A test covers every market by default, no exceptions, unless the market's requirements say otherwise. It's a small flip on paper, but it changes the whole posture of the suite: instead of proving a test works in each market one by one, you now only have to call out where it doesn't.

Honestly, a refactor like this used to take months too, the kind of thing you'd quote a quarter for. With AI helping churn through the repetitive parts, we got it down to weeks.
Markets aren't identical under the hood: flows like onboarding (like the snippet above) and payment methods differ. That doesn't disappear with opt-out-by-default, it's just handled differently. A market that does something different gets an explicit tag, like @skip-dk above, instead of a test simply never existing for it.
A few other design decisions were part of the plan from the start:
Locators got simpler for one. Instead of a mix of raw CSS selectors (like '.wallet-btn' typed directly into a test), test IDs, and one-off helper functions living inside classes, a locator is now just a function that returns the element you want.

No inheritance chain to trace, nothing to extend. For values that differ per market, such as button copy or labels, there's a tiny helper that resolves the right value instead of every test branching on market itself:

Nothing fancy. But "this one thing is different in Denmark" now lives in one small place, not scattered across the suite as a number of different if (market === 'dk') checks that somebody has to remember exist.
“Make the change easy, then make the easy change. ”Kent Beck
Setting up the test state also stopped going through the UI as much. Clicking through multiple screens just to get a test into the right starting state is slow, and worse, it's flaky, every click along the way is a chance for something completely unrelated to break. So where we can, the state just gets set up directly and the UI is left for tests that are actually about the UI.

One call to a dedicated test-helper service, and the test starts with a fully set-up user, already verified and ready to go, instead of burning the first third of the test just getting there through the UI.
Failures fail fast too. Before, if something upstream was down, we'd find out the hard way, multiple tests timing out one after another over the better part of an hour while CI budget quietly burns. Now there's a quick check before anything even starts:

Something's broken? You know in about 5 seconds, with the actual reason, not half an hour of ❌'s that all secretly mean the same thing.
Moreover, onboarding a new market got boring (in the best way possible). It used to mean copying and adjusting a pile of existing tests. Now it's basically adding one line of config and the rest of the suite already covers you, because "works everywhere" was the default the whole time.
One little DevEx win: running tests locally used to mean remembering the exact combination of env vars and flags every time, or more honestly, spamming the "Up arrow" key on the command line, doomscrolling through your shell history to find that one familiar command you used last week. So there's now a tiny interactive command that just asks you what you need. Pick the market, environment, test suite, hit enter, go. Sounds minimal, but it's the kind of small thing that decides whether people actually run tests locally or just push and pray.
Surprisingly, adoption wasn't the hard part; once people saw a market and functionality onboard in an afternoon instead of a sprint, the rest sold itself. We started with two markets and duplicated tests, one copy per market, slowly drifting apart. Now we run one test across n markets by default, no duplication. Onboarding a new market is a config change, not a project one.
It's on a solid foundation now. And the best part isn't the framework itself, it's watching folks pick it up and start contributing tests on their own with ease and speed, confidently, against multiple markets. No wading through legacy quirks first, no "go ask the one person who understands the old setup." That's the real payoff of groundwork like this: it becomes everyone's framework.