Write a sentence describing the prediction moment: for example, the model scores a new case using only information available when that case arrives. Then compare every feature with that rule. An event timestamp is not necessarily the timestamp when the feature became available.
Inspect the split unit. Repeated records from one person, customer or device can make a random row split much easier than production. A time split can still leak if a delayed label or aggregated feature includes observations from the future. Choose groups, time boundaries and label maturity according to the actual task.
Follow every fitted transformation. Scaling, imputation, encoding, feature selection and threshold choice should learn from the permitted development data. A pipeline fitted within each training fold helps, but it cannot correct a feature that already contains future information.
Finally inspect the human loop. Repeatedly choosing changes because they improve the test score adapts the project to that test set. A credible final report describes the untouched holdout, selection procedure and remaining uncertainty. A lower but valid estimate is a stronger basis for a deployment decision than an impressive score from an invalid boundary.