Week two of the AWSMinds Agentic Football Cup ended 6-5. Our win, decided in the third minute of sudden death, by a midfielder named Petit scoring his fourth goal of the match. If you only read the scoreline: greatest week of the season. If you watched the games: my team spent half the week frozen by its own instructions, and the other half pressing so hard the opponents scored on the counter.
This is the follow-up to week one, where my AI football team taught me that a slow model is a statue and a vague rule is a player standing still. Week two taught me something more uncomfortable: the biggest enemy on the pitch is your own prompt.
The 6-5 match was the best and worst game of the season
We beat Dusk Mustangs in sudden death, with three lead changes, and six of our seven shots on target. It was also the week I discovered my own formation had quietly stopped existing.
The culprit was a rule I was proud of. My midfielders were told to stay behind the striker — sensible defensive shape, textbook stuff. Except the engine reads "behind" literally. When a midfielder found himself ahead of the striker, he dutifully turned around and jogged backwards. Both of them. All match. We attacked with one man while two midfielders patrolled backwards like security guards at a concert that had already ended.

The fix is one sentence in a prompt, but you only find it by watching your midfielder walk the wrong way for two minutes of game time. The rule now: any positional anchor needs an explicit exception for when the player is already on the far side. Small models do not interpret. They obey, in the most literal way possible.
The press storm: when your own defense eats your team
Later in the week we lost 1-6 to a team that does nothing but defend in a wall. Our agents pressed them 313 times — 57% of all commands — against an opponent whose entire strategy was to be pressed. Every loose ball spawned three of our chasers, our shape vanished, and their three defenders scored more goals than their entire attack.
Nothing was broken. That's what made it interesting. Two of my rules — "closest player presses the free ball" and "press the ball carrier in their half" — were each sensible, and together they generated a swamp. The fix took one line: when three or more opponents cluster within four steps, that's a wall, and walls are not pressed, they are walked around.
In the next match our press count went from 313 to 7, and we won 4-2. Same players, same opponents' style, one deleted instruction.

The leaderboard math that changed everything
Then I sat down with the points table — and got it wrong first. That week's ledger fit perfectly with a model where a clean sheet pays 15 points, five wins' worth. It was a coincidence of arithmetic: one week of results can fit a wrong formula exactly. The official rules say a win is +30, a draw +10, a loss +2. Bonuses ride on wins only: +6 per goal of margin, capped at +30, +10 for a clean sheet, and +8 per win once a streak reaches three. A clean sheet is a third of a win, not five of them.
The corrected table changed the strategy anyway. Blowouts are worth banking — a 5-0 win pays 30 + 30 + 10, more than triple a narrow one, so when you can rout a team, rout them. Streaks reward sequencing: stack wins, don't spend them on experiments. And until 27 matches are played, every single result counts in full toward the season ranking. There is no throwaway match.

So week three has a new philosophy. Instead of "how do we score more", the whole prompt book is being rebuilt around one question: how do we never get scored on? Concretely, this week's big change: my goalkeeper is no longer allowed to pass to anyone standing near him. He waits, with the ball, until his teammates push upfield, and only then does the ball leave his box. A keeper holding the ball for a few seconds feels wrong. A keeper passing it to a defender with an opponent two steps away costs 10 points.
The coach is also a model — and also makes things up
The platform sends post-match recommendations from its own AI coach. This week it confidently suggested "rebalance your pressing" based on a match in which the telemetry claimed we pressed 378 times in two minutes at 174-millisecond latency. I cross-checked: the physical ceiling of my rule set is about 127 presses, and that latency is below what the model can physically produce. The stats were glitched, the recommendation was fiction, and my players had actually played a normal, disciplined match.
I still took the coach's advice under review, because that's the job — but a recommendation built on impossible data gets declined, not implemented. Trustworthy automation is exactly this: the system proposes, you verify the ground truth, and you keep the veto. Whether the system is booking meetings or rewriting your football formation, someone has to own the part where reality gets checked.
That's what we build for clients: automation with a veto. If your business has a process that keeps making confident recommendations based on data nobody audits, that's worth a conversation — start here, or bring us the process and we'll tell you which parts to automate and which parts to keep human.
Week two ledger: five wins, four losses, 21 goals scored, 25 conceded, and a 45-place climb up a table of 427 teams. Week three is the clean-sheet project. Zico, the goalkeeper with the hat-trick from week one, is still out there — and now we know exactly what his goals are worth.
