Reported October 2026
Capital One

Predict Taxi Driver Classes

Reported by candidates from Capital One's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Capital One OA. Under 2s to a working solution.
Founder's read

The mistake that sinks a first attempt on this Capital One OA from October 2026 is treating it like a SQL query. It isn't. You get train_data, validation_data and test_data tables, and you have to train a binary classifier in the Pandas runtime, then output one column named driver_class for every test row. Class 1 (B) is the positive class. Don't feed driver_id into the model. If you blank on the setup, StealthCoder runs invisibly as a safety net during the live OA and gives you a working pipeline. Know the shape of the answer before you open the editor.

The problem

You are given preprocessed train_data, validation_data, and test_data tables. Train a binary classifier using the labeled training and validation rows, then predict driver_class for every test row.
0 represents A class.
1 represents B class and is the positive class.
Do not use driver_id as a model feature.
You may choose any deterministic classification method available in the Pandas runtime.
Return exactly one column named driver_class, with predictions in the original test_data row order. The objective is to identify all positive rows while avoiding false positives; each published fixture has one exact expected label per test row.

Tables
train_data: driver_id PK (Integer), car_model (Integer), car_manufacture_year (Integer), days_since_inspection (Integer), age (Integer), experience (Integer), second_language (Integer), rating (Decimal), net_worth_of_tips (Decimal), number_of_rejected_rides (Integer), number_of_upvotes (Integer), number_of_complaints (Integer), number_of_incidents (Integer), driver_class (Integer)
validation_data: driver_id PK (Integer), car_model (Integer), car_manufacture_year (Integer), days_since_inspection (Integer), age (Integer), experience (Integer), second_language (Integer), rating (Decimal), net_worth_of_tips (Decimal), number_of_rejected_rides (Integer), number_of_upvotes (Integer), number_of_complaints (Integer), number_of_incidents (Integer), driver_class (Integer)
test_data: driver_id PK (Integer), car_model (Integer), car_manufacture_year (Integer), days_since_inspection (Integer), age (Integer), experience (Integer), second_language (Integer), rating (Decimal), net_worth_of_tips (Decimal), number_of_rejected_rides (Integer), number_of_upvotes (Integer), number_of_complaints (Integer), number_of_incidents (Integer), driver_class (Integer)

Constraints
Each input table is non-empty.
Training and validation labels are 0 or 1, and both classes occur in the combined labeled data.
Test labels are null and must not be used.
Every feature value is finite.
All three tables have the same feature columns.
The output has exactly one row per test row.

Reported by candidates. Source: FastPrep

Pattern and pitfall

This is a supervised classification task dressed up as a table problem. The plan is short. Concatenate train and validation, drop driver_id and driver_class from the features, fit a deterministic model such as logistic regression or a random forest with a fixed random_state, then predict on test_data. The pitfalls are concrete. Reordering rows breaks the output, so keep the original test_data order and don't sort by driver_id. Using driver_id as a feature leaks noise. Leaving the seed unset makes predictions non-deterministic, and the fixtures expect one exact label per row. The goal is to catch every positive while avoiding false positives, so check precision and recall on validation before you refit on everything. Don't use the null test labels. Return a single column named driver_class and nothing else. If the pipeline stalls mid-assessment, StealthCoder is the hedge that reads the prompt and hands you the fit and predict code.

StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.

If this hits your live OA

You can drill Predict Taxi Driver Classes cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Capital One's OA.

Capital One reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Predict Taxi Driver Classes FAQ

How hard is Predict Taxi Driver Classes really?+

The code is easy, the setup is where people slip. You need a clean feature list, a fixed seed, and the right output shape. If you've used scikit-learn or any Pandas-compatible classifier, it's a short script. The trap is the format, not the math.

What's the trick to getting the labels right?+

Fit on train plus validation combined, exclude driver_id and driver_class from the features, and use a deterministic model with a fixed random_state. Predict on test_data in its original order. Fixtures expect one exact label per row, so reproducibility matters more than a fancy model.

Should I use driver_id as a feature?+

No. The prompt says not to. It's an identifier, so it carries no real signal and can hurt generalization. Drop it from the feature matrix before fitting. Keep it only if you need to line rows up, and even then rely on the original test_data order.

What should the output look like?+

Exactly one column named driver_class, with one row per test row, in the original test_data order. Values are 0 or 1. Don't add driver_id, probabilities, or an index column. Extra columns or reordered rows will fail the fixture comparison.

How do I prepare for this in 48 hours?+

Write a small scikit-learn pipeline from memory: load, split features and target, fit logistic regression or a random forest with a fixed seed, predict, return a one-column DataFrame. Practice checking precision and recall on a validation split. Then run it once end to end on any tabular dataset.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Capital One.

OA at Capital One?
Invisible during screen share
Get it