Suppose the car analysis will run each month on 100,000 new rows with the same fields. The team must calculate the same group counts, inspect rows behind its claims, and write the report wording.
Again, I'm not sure why you would be running this check multiple times. If you always reported accuracy in the past n months of the cars with all the data, you'd better need a really strong prior to believe that you should filter out these exact training data matches, otherwise you're just p-hacking by doing multiple hypothesis testing (like treating each month as as new hypothesis for changing your reporting)