STEM Thinking Skills Assessment

Scientific Thinking

6 minCourse moduleTopic 2.6

§ 01

Overview

Module 2.6: Scientific Thinking

Thinking Like a Scientist

The STEM Thinking Skills Assessment doesn't test whether you've memorized the periodic table or can name every bone in the human body. It tests whether you can think like a scientist --- which means you can look at data, form reasonable explanations, design fair experiments, spot flaws in reasoning, and draw defensible conclusions.

This is great news. Scientific thinking is a skill, not a knowledge base. You can practice it and get better at it regardless of what science courses you've taken. This module walks you through each component of scientific reasoning and gives you real scenarios to practice with.

Hypothesis Generation from Given Data

A hypothesis is an educated guess that explains an observation --- but it's not just any guess. A good hypothesis is specific, testable, and based on the evidence in front of you.

What Makes a Hypothesis Strong?

  • It explains the data you've observed

  • It makes a prediction you can test

  • It could potentially be proven wrong (this is called "falsifiable")

  • It doesn't overreach beyond what the data supports

Example: The Wilting Plant

You notice that the plants on the left side of a greenhouse are wilting, but the plants on the right side look healthy. Both sides get the same amount of water. What might explain this?

Weak hypothesis: "The left plants are dying because something is wrong."

This is too vague. What's wrong? How would you test it?

Better hypothesis: "The left side plants are wilting because they receive more direct sunlight through the west-facing windows, causing them to lose water faster through transpiration."

This is specific, explains the observation, and you can test it --- measure the sunlight on each side, check soil moisture levels, or move some plants from left to right to see if they recover.

Practice: Generate Your Own

A teacher notices that students who sit in the front row of her class score an average of 8 points higher on tests than students who sit in the back row.

Try generating at least three different hypotheses that could explain this:

  1. Students who choose to sit in front may be more motivated and engaged, leading to better attention and higher scores. (The seating doesn't cause better grades --- motivation causes both.)

  2. Students in the front can see the board more clearly and hear the teacher better, giving them better access to the material. (Physical proximity improves learning.)

  3. The teacher may unconsciously give more attention, make more eye contact, and direct more questions to front-row students, increasing their engagement. (Teacher behavior, not student choice, drives the difference.)

Notice how each hypothesis suggests a different cause and would require a different experiment to test. That's the kind of thinking the STEM assessment is looking for.

Experimental Design

Understanding how to design a fair experiment is one of the most important scientific thinking skills. The key concepts are variables --- the things that can change in an experiment.

The Three Types of Variables

Independent variable --- The thing YOU deliberately change. You choose it. You control it. There should be only one independent variable in a well-designed experiment.

Dependent variable --- The thing you MEASURE to see if your change had an effect. It "depends" on what you did to the independent variable.

Controlled variables (constants) --- Everything else that you keep the SAME so the experiment is fair. If you change multiple things at once, you won't know which change caused the result.

The Memory Trick

The independent variable is what I change.

The dependent variable is what I measure (the data I collect).

The controlled variables are what I keep the same.

Example: Does Music Affect Study Performance?

Hypothesis: Students who study with classical music score higher on vocabulary tests than students who study in silence.

  • Independent variable: Background sound (classical music vs. silence)

  • Dependent variable: Vocabulary test scores

  • Controlled variables: Same vocabulary list, same amount of study time, same test, same room temperature, same time of day, similar group sizes

Why do we need controlled variables? Imagine you let the music group study for 30 minutes and the silent group study for 15 minutes. If the music group scores higher, was it the music or the extra study time? You can't tell. That's why everything except the one variable you're testing must stay the same.

Control Groups

A control group receives no treatment --- it's your baseline for comparison. In the music experiment, the group studying in silence is the control group. Without it, you'd have no way to know if classical music helped, hurt, or did nothing.

Sample Size Matters

Testing one student with music and one without doesn't tell you much --- maybe that one student is just better at vocabulary. The more students you test, the more confident you can be that your results reflect a real effect and not just random variation.

Practice: Design the Experiment

A student wants to know: Does the color of light affect how fast bean plants grow?

Design the experiment:

  • Independent variable: Color of light (red, blue, green, white)

  • Dependent variable: Plant height measured in centimeters after 14 days

  • Controlled variables: Same type of bean seeds, same amount of soil, same pot size, same amount of water, same temperature, same number of hours of light per day

  • Control group: Plants grown under regular white light

  • Sample size: At least 5 plants per light color (to account for natural variation between individual plants)

Data Analysis

On the STEM assessment, you'll encounter data presented in tables, graphs, and charts. Your job is to read them accurately, identify trends, and draw reasonable conclusions.

Reading Tables

A table of experimental results:

A researcher measured how temperature affects the time it takes for sugar to dissolve in water (using 10g of sugar in 200mL of water each time):

Temperature 20°C: 120 seconds

Temperature 40°C: 78 seconds

Temperature 60°C: 45 seconds

Temperature 80°C: 22 seconds

What's the trend? As temperature increases, dissolving time decreases. The relationship is inverse --- when one goes up, the other goes down.

Is the decrease constant? From 20 to 40 (a 20-degree increase), time dropped by 42 seconds. From 40 to 60, it dropped by 33 seconds. From 60 to 80, it dropped by 23 seconds. The rate of decrease is slowing down --- each additional 20-degree increase has less effect. This suggests the relationship is not linear.

Reading Graphs

When you see a graph, ask yourself:

  1. What's on the x-axis (horizontal)? This is usually the independent variable.

  2. What's on the y-axis (vertical)? This is usually the dependent variable.

  3. What's the overall trend? Going up? Going down? Flat? Curved?

  4. Are there any unusual points that don't fit the pattern (outliers)?

  5. What does the title tell you about context?

Common Graph Pitfalls

  • Scale tricks: A graph might start the y-axis at 50 instead of 0, making small differences look enormous. Always check the scale.

  • Correlation isn't causation: Two things moving together on a graph doesn't mean one causes the other. Ice cream sales and drowning rates both go up in summer --- but ice cream doesn't cause drowning. Hot weather drives both.

  • Missing context: A graph showing "sales went up 200%" sounds impressive until you realize sales went from 1 unit to 3 units.

Practice: Analyze This Data

A class tested different amounts of fertilizer on tomato plants. After 30 days:

0g fertilizer: average height 15cm, 12 tomatoes

5g fertilizer: average height 22cm, 18 tomatoes

10g fertilizer: average height 28cm, 24 tomatoes

15g fertilizer: average height 30cm, 25 tomatoes

20g fertilizer: average height 26cm, 19 tomatoes

25g fertilizer: average height 20cm, 14 tomatoes

Questions to consider:

  1. At what amount of fertilizer do plants grow tallest? 15g.

  2. What happens when you add more than 15g? Growth and fruit production decrease.

  3. Why might too much fertilizer actually hurt the plants? Excess fertilizer can damage roots, change soil chemistry, or create toxic salt concentrations.

  4. What's the optimal amount based on this data? Around 10-15g --- both height and tomato production peak in this range.

  5. Is 15g clearly better than 10g? The height difference is only 2cm and the tomato difference is only 1. With natural variation between plants, this difference might not be meaningful. You'd need to run the experiment again (replicate it) to be confident.

Drawing Conclusions from Experimental Results

Drawing a conclusion means connecting your results back to your hypothesis. Did the data support it or not?

Good Conclusions Are:

  • Directly supported by the data (not by what you wish the data showed)

  • Specific about what the evidence shows

  • Honest about limitations

  • Careful not to overgeneralize

Example of Overgeneralizing

Experiment: You tested whether playing brain-training games for 20 minutes a day improved math quiz scores among 25 eighth-graders at one school over 4 weeks.

Result: The group that played brain-training games scored an average of 6 points higher on math quizzes.

Overgeneralized conclusion: "Brain-training games make all students better at math."

Better conclusion: "In this study, 8th-grade students who played brain-training games for 20 minutes daily for 4 weeks scored an average of 6 points higher on math quizzes than students who did not. Further research with larger groups and longer time periods would help determine whether this effect is consistent and lasting."

The second conclusion is careful. It specifies who was tested, for how long, and what was measured. It acknowledges that one study isn't the final word.

Evaluating Validity of Scientific Claims

The STEM assessment may present you with a scientific claim and ask you to evaluate whether it's well-supported. Here's what to look for:

Red Flags That Weaken a Claim

  • Small sample size: "We tested 3 students and found..." --- three people is not enough to draw general conclusions.

  • No control group: Without a comparison group, you can't know if the treatment caused the result or if it would have happened anyway.

  • Correlation claimed as causation: "Students who eat breakfast get better grades, so eating breakfast causes better grades." Maybe --- but maybe families that prioritize breakfast also prioritize homework, or students who are less stressed eat more regularly.

  • Cherry-picked data: Presenting only the results that support your point while ignoring results that don't.

  • Vague or unmeasurable claims: "This supplement boosts brain power" --- how do you measure "brain power"?

  • No replication: A result that happened once might be a fluke. Science gains confidence through repeated results.

Practice: Evaluate These Claims

Claim 1: "A study of 500 high school students found that those who slept 8+ hours per night had GPAs 0.4 points higher than those who slept fewer than 6 hours per night."

Evaluation: Decent sample size (500). But this is observational, not experimental --- students weren't randomly assigned to sleep amounts. Students with higher GPAs might have better time management, leading them to both finish homework earlier and get more sleep. The sleep-GPA connection is a correlation, and this study alone can't prove causation.

Claim 2: "My cousin drank green tea every day and never got sick all winter. Green tea prevents colds."

Evaluation: This is anecdotal evidence --- a single person's experience. There's no control group (what if your cousin also washes hands frequently, avoids crowded places, or just has a strong immune system?). Sample size of one. No controlled variables. This claim is not scientifically supported.

Claim 3: "Researchers randomly assigned 200 students to two groups. One group used a new study app for 6 weeks. The other group studied the same material using traditional methods for the same amount of time. The app group scored 12% higher on a standardized test."

Evaluation: Strongest of the three. Random assignment helps ensure the groups are comparable. There's a control group. The sample size is reasonable. The same material and same time period are controlled. This is a well-designed experiment. The one question you might ask: was it replicated? One study is a good start, but replication builds confidence.

Identifying Flaws in Experimental Methodology

This is where you put it all together. The STEM assessment loves questions that describe an experiment and ask you to identify what's wrong with it.

Common Experimental Flaws

Confounding variables --- When more than one thing changes between groups, you can't tell which change caused the result.

Example flaw: A school gives one class a new textbook AND a new teacher, while another class keeps the old textbook and old teacher. If the first class performs better, was it the textbook or the teacher?

Selection bias --- When the groups being compared aren't truly comparable.

Example flaw: A study compares students who voluntarily joined a tutoring program with students who didn't. The volunteers are probably already more motivated --- so any improvement might be due to motivation, not tutoring.

Measurement problems --- When the way you measure the outcome is flawed.

Example flaw: A study on whether exercise improves mood asks participants to self-report their mood, but participants know the researchers expect exercise to help. They might report feeling better just because they think they should (this is called response bias).

No baseline measurement --- If you don't measure before and after, you don't know what changed.

Example flaw: A school implements a new anti-bullying program and finds that 15% of students report being bullied at the end of the year. Is that good or bad? Without knowing the rate before the program started, you can't tell if it went up, down, or stayed the same.

Practice: Spot the Flaw

Experiment: A student wants to test whether plants grow better with tap water or bottled water. She plants one seed in a pot on her windowsill and waters it with tap water. She plants another seed in a different-sized pot in her basement and waters it with bottled water. After three weeks, the tap water plant grew taller.

How many flaws can you spot?

  1. Different pot sizes --- this changes how much room the roots have to grow (confounding variable)

  2. Different locations --- the windowsill gets sunlight, the basement may not (confounding variable)

  3. Only one plant per group --- no way to account for natural seed variation (sample size too small)

  4. She likely used different soil amounts because of different pot sizes (another confounding variable)

  5. Temperature, humidity, and air circulation probably differ between a windowsill and a basement (more confounding variables)

This experiment has so many uncontrolled variables that the results are essentially meaningless. A better design: same pot size, same location, same soil, same amount of water, multiple plants per group --- the only difference being tap versus bottled water.

Putting It All Together: A Full Scenario

Read this mini-experiment and answer the questions that follow:

A student hypothesizes that listening to nature sounds (rain, birdsong, ocean waves) while taking a math test will improve test scores. She recruits 40 classmates and randomly divides them into two groups of 20. Group A takes a 30-question math test in a quiet room. Group B takes the same test in the same room at a different time, with nature sounds playing through speakers. Both groups are given 45 minutes.

Results: Group A average score: 22.3 out of 30. Group B average score: 23.1 out of 30.

  1. What is the independent variable? The presence or absence of nature sounds.

  2. What is the dependent variable? Math test scores.

  3. Name three controlled variables. Same test, same room, same time limit. (Also: same number of students per group, random assignment.)

  4. Is this a well-designed experiment? Mostly yes --- random assignment, control group, same test, adequate sample size. One concern: the groups tested at different times, which could introduce variation (morning vs. afternoon alertness).

  5. Do the results strongly support the hypothesis? The difference is less than 1 point (22.3 vs. 23.1) on a 30-point test. This is a very small difference that could easily be due to random chance rather than the nature sounds. You'd need to run a statistical test to determine if this difference is significant. Based on the raw numbers alone, the evidence is weak.

  6. What would make this study more convincing? Larger sample size, testing at the same time of day (or randomly alternating which group tests first), repeating the experiment multiple times, and running a statistical significance test on the results.

Key Takeaways for Test Day

  • Scientific thinking questions reward careful, logical reasoning --- not memorized facts.

  • When asked to identify variables, remember: independent is what changes, dependent is what you measure, controlled is what stays the same.

  • Always look at experimental design with a critical eye. Ask: what else could explain these results?

  • Be skeptical of strong conclusions from weak evidence. The best scientific thinkers are the most cautious about overgeneralizing.

  • When generating hypotheses, make sure they're specific, testable, and grounded in the data you've been given.

  • Practice reading data carefully. The answer is in the numbers --- not in what you expect or hope the numbers show.