A train test split is the cut you choose before you write the fitting code. Random when the rows are independent and the classes are roughly even. Stratified when one class is rare. By time when later rows must stay unseen. One of those three, for the table you actually have. A random split on the wrong table produces a score that will not survive the next file.
I have always written software. For more than five years that work was data in Informatica: mappings, extracts, and the load that had to land clean. The MSc in applied machine learning came after, and it gave names to a failure I had already refused on a bad extract. If the rows you score were already inside the fit, the number at the bottom of the notebook is a replay. It is not a test of the next load.
The helper will not make this choice for you. train_test_split shuffles unless you stop it. Shuffle is a random split. It is the right cut only for the first case above. On a rare class, or on a log that arrives in order, the default runs, the printouts look tidy, and the score dies the moment a new file shows up. Pick the cut from the table. Then write the few lines that implement that cut, and leave the held-out rows untouched.
What a train test split is
You take one table and divide it into two sets of rows. The training rows are the only rows the algorithm may see while it fits. The test rows are held out. Their labels are for the score you compute after the fit, not for the fit itself. Every row goes to one side. You do not score a row that also sat in the training matrix.
The cut is by rows, not by columns. A column you will have at prediction time can appear on both sides. The label column appears on both sides too, with a different job on each side. A column you will not have when a new row arrives does not belong on either side. On the telecom churn table I drop the customer id before the split. An id is a key. It is not a reason a customer leaves. Leaving it in teaches the model to recognize a key it will never see again in that form, or, worse, a key that repeats.
A split is also not a second copy of the same rows. If a merge left the same business key in the table twice, a random cut can send one copy to train and one copy to test. I have seen that pattern in extracts that were joined on a sloppy key: the test row is the training row with a different index. Deduplicate on the business key before you cut, or the score is a memory test with extra steps.
People spend their time on the fraction. On the logistic churn model I used 70 percent for train and 30 percent for test. In the later beginner tutorial I used 80 percent and 20 percent. Both fractions are fine. Neither fraction corrects a bad kind of cut. Twenty percent of a shuffled log is still a shuffled log. Thirty percent of a rare class, drawn at random from a short table, can still miss that class. Decide the kind of split first. The percentage is the smaller knob, and you can change it without changing the meaning of the test.
Set random_state when you shuffle, so the same notebook produces the same cut tomorrow. That is reproducibility. It is not a substitute for the right cut. A fixed seed on a shuffled time series is a repeatable mistake.
Why those rows have to stay out
A fitted model has already used the training rows. A score on those rows answers a weak question: can the procedure repeat what it just saw? I have watched a logistic regression look settled on the training fold and then miss the pattern that mattered on rows from a later extract. The training number did not warn me. The extract did.
You hold rows out so the score refers to rows the fit did not touch. That is the number I will put in front of someone who might act on it. In the load work, the next thing that arrives is a new file. It is not a reshuffle of the file I just mapped. If the test rows were drawn from a shuffle of the current file, I have measured the shuffle. I have not measured the next file.
There is a second use of the seal. Anything you change while watching the test score becomes part of the fit. If I drop a column because the test score rose, those test rows helped choose the model. The number I quote after that drop is no longer a test. When I need to compare two settings, I cut a validation slice from the training rows only, or I cross-validate inside the training rows. The beginner tutorial’s cross-validation step is that inner check. It does not replace a final set of rows you have not used for any choice. Touch the test rows when the model is frozen, compute the score, and stop tuning.
Write down, in one line next to the score, which cut you used and what question it answers. “Random, one row per customer, classes near even” is a different claim from “last six weeks, no shuffle.” Readers of your notebook, including you in a month, should not have to recover that from a seed.
Random, when the rows are independent
A random split shuffles the rows, then takes a fraction for the test side. That is what train_test_split does if you leave shuffle at its default, which is true. Each row has the same chance to land in the test set. Nothing about the order of the file survives.
Use it when two facts are true at once. The rows are independent, and the classes are roughly even. Independent means one row does not carry the identity of another row that might fall on the other side. Each customer appears once, or each order appears once, and there is no clock you are obliged to respect. Roughly even means a random slice still contains both classes in about the same mix as the full table. A yes and a no that both show up in large numbers. A regression target with no rare bin you must protect. For a plain regression on independent rows, there is no class to preserve, and a random cut is the usual start.
The appeal is real. The code is one call. The test rows are a mixture of the same population as the training rows, which is what you want when the next row really is “another draw like these.” A software test of a pure function has the same shape: inputs that are not ordered, and a holdout that could have been any input. I trust a random split in that setting.
It fails in three ways I have actually had to unwind.
The rare class comes first. Fraud, a defect, a minority churn flag: a small share of the rows. A random 20 percent of a large table usually still contains some of them, but the rate wobbles, and on a short table the test side can contain none. Precision and recall on a class that is absent are not a property of the model. They are a property of an unlucky draw. I count the classes on both sides before I fit. If the test side is missing the class I care about, I do not “run it anyway.” I change the cut.
The repeated key comes second. Same customer, same account, same device, several rows, split at random across the two sides. The model can memorize the customer. The score reads like generalization. It is recognition of an id, or of a bundle of features that only that id has. Stratifying the label does not fix this either. The unit of the split has to match the unit you will be judged on. If that unit is the customer, keep every row of a customer on one side. GroupShuffleSplit does that when the key repeats and you are not forecasting a later period. A random split will not.
Time comes third. A log, a daily snapshot, a sequence of loads: shuffle puts later rows on the training side and earlier rows on the test side. The model is allowed to see Thursday while it is scored on Tuesday. The next file will not arrive in that order. The score will not survive it. If a time column is the time of the decision, and later rows must stay unseen, you are not in the random case, however even the classes look.
Stratified, when one class is rare
A stratified split draws inside each class, then combines the draws. If 8 percent of the table is the positive class, about 8 percent of the training rows and about 8 percent of the test rows are positive. The rate is preserved. The rows within each class are still shuffled.
In sklearn you pass the label as stratify=y. The argument is the class column, not a feature, and not a row number. test_size still sets the fraction. random_state still fixes the draw. What changes is that each class is represented on both sides in proportion to its share of the table.
Use this when one class is rare and the rows are still independent of each other and of any clock you must respect. Churn is this case. It is the smaller class. In the beginner tutorial I pass stratify=y on that label so the test slice keeps the churn rate. On the earlier logistic model I called train_test_split with a 70/30 random cut and no stratify. That default is the wrong one once the class you care about is the smaller one. The tutorial is the call I would copy.
Check the counts after the call, not only the shapes. Shapes tell you that 20 percent of the rows moved. Counts tell you that the rare class moved with them. If you print one number, print the positive count on each side.
Stratified fails when you needed a clock. It still shuffles. A time-ordered table with a rare event is not repaired by stratify=y. You cut by time, then you look at the future window. If that window contains none of the rare class, the metric you wanted cannot be computed there. Widen the window until the event appears, or accept that this particular tail cannot score that class. Do not shuffle the table to force the counts to look balanced. You would be buying a prettier rate with a leaked future.
It also fails when the unit is a group. Two visits from the same patient can land on opposite sides of a stratified split, because the split balances the label and ignores the patient id. If your real problem is a repeated key and there is no forecast horizon, group the split. Stratify will look correct in the class rates and still leave the same patient on both sides.
And it fails when you stratify on the wrong column. The column has to be the outcome you will score. Stratifying on a feature forces that feature’s mix to match across the cut, and it can hide a shift you needed to see. If the “label” was written by a process that already knew the outcome you are trying to predict, matching its rate on both sides preserves a leak. I treat that as a column problem first: if you would not have the field at decision time, it is not a label you may stratify on, and it is not a feature you may keep.
For a continuous target there is no class, so stratify has nothing honest to receive unless you bin the target yourself. I rarely do that. If the rows are independent, I use a random split for regression. If a thin tail of the target is the whole point of the model, say a rare loss band, bin only for the split, train on the original target, and still check that the tail appears on the test side. If the rows are a series, I ignore bins and cut by time.
By time, when later rows must stay unseen
A time split sorts by the moment the decision was made, then cuts. Early rows train. Later rows test. Nothing is shuffled. In code this is a slice after a sort, not a helper that mixes rows. I sort on the event time, take the first part as train, and hold the last part as test. Then I assert that the latest training timestamp is strictly earlier than the earliest test timestamp. If that assert fails, the column I trusted is not an order.
Use it when the question is about what comes next. The next day’s load, a sensor log, a claims file that arrives in sequence, a forecast. The question is: if I had stopped at this date, would the fit have worked on what came after? A random split cannot ask that. It destroys “after.” A stratified split cannot ask it either, because the rare class is still drawn from the whole clock.
On mapping work this is the cut I insist on whenever the extract is a sequence. The file has an order even when nobody labeled it as a time-series project. A snapshot date, a transaction time, an event time: if later rows must stay unseen, they stay on the test side, all of them, and the training side stops at the cut. Mixing in a few “helpful” rows from after the cut is how a demo score gets built. It is also how that score misses the next file.
The same customer may appear before the cut and after it. For a forecast, that is often legitimate. You are allowed to know the customer’s past. You are not allowed to know the customer’s future. A time split keeps the future row’s label out of the fit. A random split does not, and it can train on a later visit while scoring an earlier one. So “this id is on both sides” condemns a random split. It does not, by itself, condemn a time split. What condemns the time split is a feature on the future row that could only have been known after the outcome, or a timestamp that is not the decision time.
That wrong clock is the failure I watch for. A batch timestamp stamped on every row when the file landed is the same value down the column, or it ties a whole file to the hour the job ran. Splitting on it does not recreate the order of events. A last-updated column that moves after the label is known is worse: the time itself has seen the outcome. The cut has to use the time you would have had when you made the prediction. If you cannot point to that column, you do not have a time split yet. You have a sort on a convenient field.
The default shuffle will quietly undo the sort. If you do use train_test_split for a time cut, the frame has to be sorted already, and you have to pass shuffle=False. I prefer the slice. There is no default left that can reshuffle it.
A time cut also answers a narrower question than people hear. The last weeks of December test December. They are a weak test of April. Say which window you scored. If the business cares about every season, one tail is not the whole claim, and you may need more than one cut date, each still keeping its own future unseen. That is still a time rule. It is not a license to shuffle.
Finally, check the rare class inside the future window. A clean time cut can still leave you with a test period that contains none of the events you care about. The order is right, and the metric is empty. Widen the test window from the recent end, or pick a cut date that leaves events on the test side, without pulling any of those test rows back into training.
A short sklearn split you can run
The block below builds all three cuts on a synthetic table. It does not fit a model. These rows are noise from a generator, so the prints are sizes and class counts. Read those counts, and read the assert. The assert is the time split checking itself.
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(42)
n = 2000
X = rng.normal(size=(n, 3))
# Independent rows, classes roughly even: random split (shuffle is the default).
y_even = rng.integers(0, 2, size=n)
X_train, X_test, y_train, y_test = train_test_split(
X, y_even, test_size=0.2, random_state=42
)
# Rare class, rows still independent: keep the class rate on both sides.
y_rare = np.zeros(n, dtype=int)
y_rare[:100] = 1
rng.shuffle(y_rare)
X_train_s, X_test_s, y_train_s, y_test_s = train_test_split(
X,
y_rare,
test_size=0.2,
random_state=42,
stratify=y_rare,
)
# Later rows stay unseen: sort by event time, then slice. No shuffle.
frame = pd.DataFrame(
{
"event_time": pd.date_range("2024-01-01", periods=n, freq="h"),
"x1": X[:, 0],
"y": y_even,
}
)
frame = frame.sort_values("event_time").reset_index(drop=True)
cut = int(len(frame) * 0.8)
train_time = frame.iloc[:cut].copy()
test_time = frame.iloc[cut:].copy()
assert train_time["event_time"].max() < test_time["event_time"].min()
print("random sizes", len(X_train), len(X_test))
print("stratified positive counts", int(y_train_s.sum()), int(y_test_s.sum()))
print("time sizes", len(train_time), len(test_time))
On the rare label you want positive rows on both sides, in a similar proportion. On the time cut you want the assert to pass and to stay in the file. If someone later sorts and then calls train_test_split without shuffle=False, this slice is the version that does not depend on them remembering that argument.
Run one of the three paths on your real table, not all three as if they were a menu of scores. The other two are here so you can see the shape of the cut you rejected. After you choose, the test object is sealed. No scaler, no encoder, and no column picker is allowed to fit on it.
Do not fit the full table before the split
The cut is wasted if the test rows already changed the fit. The usual path is quiet. You clean the whole frame, scale the whole frame, encode the whole frame, or choose columns from the whole frame, and only then you split. The test rows helped build the transformer. The score that follows is partly a score of rows the transformer has already digested.
Do not fit a scaler, an encoder, or a feature choice on the full table before the split.
A StandardScaler learns a mean and a spread. Fit it on every row, and the test rows pull that mean toward themselves. The training data has seen a fact about the test data. An encoder, whether get_dummies or OneHotEncoder, fit on every row, can grow a column for a category that exists only on the test side. The training matrix then contains a column that exists because the future showed up. Feature choice is the same mistake in another function. SelectKBest, a correlation cutoff, or recursive feature elimination, run on every row, uses the test labels to decide which columns survive. You can split perfectly and still leak, if the selection step ignores the split.
On the logistic churn model I published, the recursive feature elimination step was fit on X and y for the whole table, even though a train and test split already existed above it. That fit ranks columns with the test labels in the room. I also drew the correlation heatmap on the full frame and dropped columns on both sides from what that picture showed. A correlation you use to discard a column is a feature choice. It has to be computed on the training rows. The split, in that notebook, was cosmetic the moment selection ran on everything. The repair is small and strict. Fit the selector on the training rows. Freeze it. Transform the test rows with the frozen selector. Do not call fit on X and y together after you have gone to the trouble of cutting them.
Cleaning and scaling belong inside the training fold only.
That includes the fill values. A median computed on the full column uses test values in the number you subtract. Compute the median on the training rows, fill the training rows with it, and fill the test rows with that same median. Do not recompute the median on the test rows. A category that never appears in training needs a defined behavior on the test side, an unknown bucket or a zero, not a second fit that adds the new category to the vocabulary.
A sklearn Pipeline is how I keep that order in the code rather than in a comment. fit on the pipeline sees the training rows. predict on the test rows applies the scaler and the model and does not refit them. The beginner tutorial builds that pipeline after the split, with the scaler inside it, and it fits feature selection on the training matrix only. That is the order to copy. Any step that learns from labels belongs in fit, on the training fold, and nowhere else.
One leak is not a scaler, and the split will not wash it out. A column that is written only after the outcome, for example a field filled in when the customer has already left, has to be dropped from the whole table before you cut. You are not learning a statistic from the test rows. You are refusing a column you will not have at decision time. The things you learn from data are the mean, the categories, the chosen columns, the filled values. Those wait until the training rows are the only rows in the room.
If you remember a single order, use this one. Drop columns you could never have at decision time. Split, with the cut this page asked you to choose. On the training rows only, clean, scale, encode, and choose features. Transform the test rows with those fitted steps. Fit the model on the training rows. Score once on the test rows.
Fine-tuning still holds rows out
If you later fine-tune a language model, hold out an eval set. You do not train on those rows. If the rows are ordered in time, you do not shuffle them. That is the same rule as the table above, and it is the whole of what this page needs to say about fine-tuning. The work here is a classical train test split in scikit-learn.
Next step
Put the cut on a real churn table, then fit only on the training side.
In Building a Logistic Regression model in Python I split a telecom table and fit logistic regression. Use that walkthrough after you have decided whether the cut should be random, stratified, or by time, and fit any column selection on the training rows only.
Or use Step 5 of the beginner machine learning tutorial. The churn split there is already stratified, and the steps after it fit the selector and the scaler on the training side. That is the pattern to keep: choose the split for the table you have, seal the test rows, and only then fit.
