Reported October 2026
Boston Consulting Group

Predict Taxi Driver Classes

Reported by candidates from Boston Consulting Group's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Boston Consulting Group OA. Under 2s to a working solution.
Founder's read

Boston Consulting Group put a taxi driver classification task in an OA reported in October 2026, and it's labeled SQL even though nothing here is a query. You get train_data, validation_data and test_data, and you have to predict driver_class for every test row. It's a binary classifier in Pandas, not a join or a window function. If you were prepping GROUP BY and self-joins, this one flips the script. The good news is the bar is simple: one deterministic model, one output column. StealthCoder sits invisibly on your screen as a safety net if you blank on the setup mid-assessment.

The problem

Use the prepared training and test data to train a classification model and produce driver-class predictions with high accuracy.
FastPrep practice interpretation
For this exercise, assume the prepared inputs are train_data, validation_data, and test_data. Train a binary classifier using the labeled training and validation rows, then predict driver_class for every test row.
0 represents A class.
1 represents B class and is the positive class.
Do not use driver_id as a model feature.
You may choose any deterministic classification method available in the Pandas runtime.
Return exactly one column named driver_class, with predictions in the original test_data row order. The practice objective is to identify all positive rows while avoiding false positives; each published fixture has one exact expected label per test row.

Tables
train_data: driver_id PK (Integer), car_model (Integer), car_manufacture_year (Integer), days_since_inspection (Integer), age (Integer), experience (Integer), second_language (Integer), rating (Decimal), net_worth_of_tips (Decimal), number_of_rejected_rides (Integer), number_of_upvotes (Integer), number_of_complaints (Integer), number_of_incidents (Integer), driver_class (Integer)
validation_data: driver_id PK (Integer), car_model (Integer), car_manufacture_year (Integer), days_since_inspection (Integer), age (Integer), experience (Integer), second_language (Integer), rating (Decimal), net_worth_of_tips (Decimal), number_of_rejected_rides (Integer), number_of_upvotes (Integer), number_of_complaints (Integer), number_of_incidents (Integer), driver_class (Integer)
test_data: driver_id PK (Integer), car_model (Integer), car_manufacture_year (Integer), days_since_inspection (Integer), age (Integer), experience (Integer), second_language (Integer), rating (Decimal), net_worth_of_tips (Decimal), number_of_rejected_rides (Integer), number_of_upvotes (Integer), number_of_complaints (Integer), number_of_incidents (Integer), driver_class (Integer)

Constraints
Each input table is non-empty.
Training and validation labels are 0 or 1, and both classes occur in the combined labeled data.
Test labels are null and must not be used.
Every feature value is finite.
All three tables have the same feature columns.
The output has exactly one row per test row.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is that this isn't an algorithm puzzle. It's a clean ML pipeline. Concatenate train and validation, drop driver_id and driver_class from the features, fit a deterministic model such as a random forest or gradient boosting with a fixed random_state, then predict on test_data. The brute-force worry doesn't apply because there's no input size pressure. The real pitfalls are leakage and format. Don't use driver_id as a feature. Don't touch the null test labels. Return a DataFrame with exactly one column named driver_class, in original test row order, so never sort or shuffle test_data. Class 1 is the positive class, so check validation precision and recall before you trust it. Accuracy alone can hide false positives. If you freeze on the pandas and sklearn boilerplate during the live OA, StealthCoder can supply it while you stay in control of the model choice.

If you see this problem in your OA tomorrow, the play is to recognize the pattern in 30 seconds. StealthCoder buys you that recognition.

If this hits your live OA

You can drill Predict Taxi Driver Classes cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Boston Consulting Group's OA.

Boston Consulting Group reuses patterns across OAs. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Predict Taxi Driver Classes FAQ

Is this really a SQL question?+

The format says SQL, but the task is a Pandas classification problem. You train a model on train_data and validation_data, then predict driver_class for test_data. Expect Python, not SELECT statements. Read the prompt carefully so you don't waste time writing queries.

What's the trick to getting high accuracy?+

Use a solid tree-based classifier with a fixed random_state, train on the combined labeled data, and drop driver_id. Check validation first to confirm precision and recall on class 1. The prompt rewards finding all positives while avoiding false positives, so don't chase accuracy alone.

What output format does it expect?+

Exactly one column named driver_class, with one prediction per test row, in the original test_data order. Values should be 0 or 1. Don't include driver_id, don't reorder rows, and don't leave nulls. Format mistakes can fail an otherwise good model.

Should I use driver_id as a feature?+

No. The prompt says explicitly not to. It's an identifier, so it adds noise or leakage. Drop it from the feature set along with driver_class before fitting, and keep it only if you need it to verify row alignment.

How do I prepare for this in 48 hours?+

Rehearse a short template: load frames, concat train and validation, split features and label, fit a deterministic model, predict, return a one-column DataFrame. Practice it once end to end so the syntax is automatic. Then spend remaining time on validation metrics for class 1.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Boston Consulting Group.

OA at Boston Consulting Group?
Invisible during screen share
Get it