import pandas as pd
check = pd.DataFrame({
"zone": ["Nord", "Sued", "Nord", "Hafen", "Nord"],
"total_eur": [10.0, 7.5, 20.0, 5.0, 30.0],
})
print(check[check["zone"] == "Nord"]["total_eur"].mean())20.0
Programming with Python
Pens down. The checkpoint is behind you. This morning the rival chain MunchCorp put out a press release bragging about their “data-driven growth.” The evidence attached: exactly one pie chart. The investor slid it across the table and said, “Yours will be better. Show me the data.”
. . .
Tobi didn’t wait. He’d pasted the question into a chatbot, proudly turned the laptop around, and hit run:
AttributeError: 'DataFrame' object has no attribute 'overview'
. . .
The AI sounded certain. The method it called does not exist. Today is about that difference, and about the tool that answers the investor for real: pandas.
Two things turn a vague request into a useful answer:
orders with columns zone (text) and total_eur (float).”total_eur for zone Nord, as a single number.”. . .
“DataFrame
orders, text columnzone, float columntotal_eur. Write one line that returns the totaltotal_eurwherezoneisNord.”
Vague in, vague out. Specific in, checkable out, and then you iterate.
Tobi’s mistake wasn’t using AI. It was shipping without checking. Every line an AI hands you gets three passes:
Five orders with values you can add in your head, and the AI’s suggested line for the Nord average:
import pandas as pd
check = pd.DataFrame({
"zone": ["Nord", "Sued", "Nord", "Hafen", "Nord"],
"total_eur": [10.0, 7.5, 20.0, 5.0, 30.0],
})
print(check[check["zone"] == "Nord"]["total_eur"].mean())20.0
. . .
The Nord orders are 10, 20 and 30. You can average those in your head. Does the AI’s answer match? Then the line has earned some trust on eighty rows. Verification is the job now, not typing.
Tobi asks the AI for the Nord average on the same five orders. It hands back one line and calls it done. What does this print?
print(check["total_eur"].mean())a) 20.0 b) an AttributeError c) 14.5
. . .
Predict first. Pick a letter, then I reveal the answer.
c) 14.5: the line averages all five orders. Nobody filtered for Nord, and pandas has no way to know that was the question:
print(check["total_eur"].mean()) # every zone
print(check[check["zone"] == "Nord"]["total_eur"].mean()) # Nord only14.5
20.0
. . .
No traceback, a plausible number, and still wrong. “It runs” only clears step 2; step 3 is where the investor’s number gets tested.
.summarize() exist?The AI wrote this for Tobi. orders is a tiny two-row frame. What happens on the last line?
import pandas as pd
orders = pd.DataFrame({"zone": ["Nord", "Sued"], "total_eur": [12.0, 9.5]})
orders.summarize()a) prints a summary table b) raises an AttributeError c) returns an empty DataFrame
. . .
Predict first. Pick a letter, then I reveal the answer.
b) AttributeError: pandas has no .summarize(). The AI invented it:
import pandas as pd
orders = pd.DataFrame({"zone": ["Nord", "Sued"], "total_eur": [12.0, 9.5]})
orders.summarize() # the method the AI inventedAttributeError: 'DataFrame' object has no attribute 'summarize'
. . .
The real method is .describe():
print(orders.describe()) # the method that actually exists total_eur
count 2.000000
mean 10.750000
std 1.767767
min 9.500000
25% 10.125000
50% 10.750000
75% 11.375000
max 12.000000
. . .
An AI that sounds sure is not the same as an API that exists.
From Session VI the course rule stands: every submission that used AI carries a one-line note saying what you used it for.
groupby line; I checked the totals by hand.”. . .
It protects you: it separates the work you understand from the work you pasted, so when a reviewer (or the investor) asks “how does this line work?”, you’re never caught claiming something you can’t explain.
Sometimes the fastest path is the one you already own. You built a whole muscle in Part I:
. . .
Use AI to draft the unfamiliar and to explain the confusing, not to skip thinking you can do yourself.
Last session, a NumPy array turned a thousand numbers into one fast object, but every value had to share one type, and columns had no names.
. . .
The convention everyone uses: import pandas as pd.
Keys become column names; each list becomes a column; one order per row, mixed types together:
import pandas as pd
df = pd.DataFrame({
"order_id": [101, 102, 103, 104, 105],
"zone": ["Nord", "Sued", "Nord", "Hafen", "Sued"],
"items": [2, 1, 3, 1, 4],
"total_eur": [18.50, 7.20, 24.00, 6.80, 31.40],
})
print(df) order_id zone items total_eur
0 101 Nord 2 18.5
1 102 Sued 1 7.2
2 103 Nord 3 24.0
3 104 Hafen 1 6.8
4 105 Sued 4 31.4
. . .
Nobody types eighty orders by hand: today’s lab loads a real CSV in one line, orders = pd.read_csv("public/orders.csv"). In the browser that line fetches over the web instead of from disk; your code doesn’t change.
.head() and .info()Before analyzing a table, glance at it. .head() shows the top rows; .info() is your first sanity check: right number of rows? Any column a surprising type?
print(df.head(3)) # first 3 rows
df.info() # columns, dtypes, non-null counts order_id zone items total_eur
0 101 Nord 2 18.5
1 102 Sued 1 7.2
2 103 Nord 3 24.0
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 5 entries, 0 to 4
Data columns (total 4 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 order_id 5 non-null int64
1 zone 5 non-null object
2 items 5 non-null int64
3 total_eur 5 non-null float64
dtypes: float64(1), int64(2), object(1)
memory usage: 292.0+ bytes
.describe().describe() computes count, mean, min, max and quartiles for every numeric column at once:
print(df.describe()) order_id items total_eur
count 5.000000 5.00000 5.000000
mean 103.000000 2.20000 17.580000
std 1.581139 1.30384 10.688873
min 101.000000 1.00000 6.800000
25% 102.000000 1.00000 7.200000
50% 103.000000 2.00000 18.500000
75% 104.000000 3.00000 24.000000
max 105.000000 4.00000 31.400000
. . .
One call, the whole numeric summary: this is the .summarize() the AI wished existed, spelled correctly.
Pick one column by name; keep rows with a boolean mask, exactly the NumPy idea from last session:
print(df["zone"]) # one column (a Series)
print()
print(df[df["zone"] == "Nord"]) # only the Nord rows, case-sensitive!0 Nord
1 Sued
2 Nord
3 Hafen
4 Sued
Name: zone, dtype: object
order_id zone items total_eur
0 101 Nord 2 18.5
2 103 Nord 3 24.0
. . .
df["zone"] == "Nord" builds a column of True/False; indexing with it keeps the True rows. .loc exists for label-based selection, but a plain mask covers today.
Tobi wants the orders over 20 EUR. He compares the column and prints the result:
big = df["total_eur"] > 20
print(big)a) a column of True/False, one per row b) only the rows whose total is over 20 c) a single True, since some order is over 20
. . .
Predict first. Pick a letter, then I reveal the answer.
a) a column of True/False. The comparison alone is the mask, a Series with one boolean per row. The rows only appear once you index with it:
big = df["total_eur"] > 20
print(big) # the mask: a Series of booleans
print(df[big]) # the rows: build the mask, then apply it0 False
1 False
2 True
3 False
4 True
Name: total_eur, dtype: bool
order_id zone items total_eur
2 103 Nord 3 24.0
4 105 Sued 4 31.4
orders.csv, eighty rows, one pd.read_csv line.head() it, ask its .shape, filter with masks, add a column on a safe copy, and answer per-zone questions with groupby. . .
Download your .py before you leave. Closing the tab without downloading loses your work, the same motion you used to hand in the checkpoint at the start of the session.
pd.DataFrame from a dict, pd.read_csv from a file; .head(), .info(), .describe() to look before you leap.df["col"], a boolean mask (df[df["zone"] == "Nord"], case-sensitive!), a new column on a .copy(), and groupby for per-category answers. pandas won’t warn you; you verify.. . .
Next episode: charts the investor can’t argue with. Numbers convince the careful; a good plot convinces the room. We turn the data room into pictures.
. . .
Working with AI this session? Revisit the AI Tools page: context-and-constraints prompting, the verify workflow, and the one-line disclosure habit. For pandas itself, the official 10 minutes to pandas guide is the friendliest next step.
. . .
For more, see the literature list of this course.