Skills Beta
Statistics and math: the one-page sheet
Mean, median, mode and range
Repeated measurements vary, so you summarize a sample with a center and a spread. The mean (sum ÷ n) uses every value and is pulled by outliers; the median (middle of the sorted data) resists them; the mode is the most common value. The range (largest − smallest) is the simplest measure of spread. Check an outlier for errors before deciding anything about it.
- Worked example. Counts 31, 28, 33, 30, 29, 35, 30, 32. Mean = 248 ÷ 8 = 31.0. Sorted: 28, 29, 30, 30, 31, 32, 33, 35; median = (30 + 31) ÷ 2 = 30.5. Mode = 30 (it appears twice). Range = 35 − 28 = 7.
- The mean uses the size of every value; the median uses only position. When a data set has an outlier or is lopsided, the two disagree, and the median describes a typical value better.
- Always pair a center with a spread. Two groups with the same mean of 25 drops can have ranges of 2 and 13: the second is far less consistent.
- Quantitative data are numbers (counted drops are discrete; measured mass or time are continuous). Qualitative data are categories. You can average numbers, not categories.
The values differ a little each time, so no single value is "the answer": you have a sample. You get the mean, a center that uses every value, for example 248 ÷ 8 = 31.0 drops. The mean is dragged toward that outlier (16.75 instead of about 14.7). The median stays with the cluster (15), so it is the better center when there is an outlier. You also report spread, starting with the range (largest − smallest), to show how consistent the measurements were.
- quantitative data
- Quantitative data are numbers from counting or measuring (discrete if counted in whole units, continuous if measured on a scale). Qualitative (categorical) data are descriptions or categories, such as color or "yes/no".
- sample (statistics)
- The set of measurements or individuals you actually collected, used to stand for the larger group you cannot measure in full. The sample size, n, is how many values it holds.
- mean (statistics)
- The arithmetic average: add every value and divide by how many values there are (x̄ = Σx / n). Every value pulls on it, including extreme ones.
- median
- The middle value once the data are sorted from smallest to largest; with an even number of values, the average of the two middle ones. Extreme values barely move it.
- mode (statistics)
- The value that appears most often in a data set. A set can have one mode, several, or none.
- range (statistics)
- The largest value minus the smallest value: the simplest measure of how spread out the data are.
- outlier
- A value far from the rest of the data. It may be a recording mistake or a real but unusual result, so it is checked and reported, not silently deleted.
- spread of data
- How much repeated measurements differ from each other. Low spread means the values cluster tightly; high spread means they scatter widely.
Graph construction
Put the independent variable on the x-axis and the dependent variable on the y-axis, each labeled with units on an even scale. Choose a line graph for a continuous independent variable, a bar graph for categories, a scatterplot for two measured variables and a histogram for a distribution. Add a key for groups. The slope (rise ÷ run, with units) describes how steeply a line changes; interpolate with care and extrapolate rarely; and remember that correlation is not causation.
- Worked example. 100 g of water: 20.0 °C at 0 min, 50.1 °C at 5 min. Slope = (50.1 − 20.0) ÷ (5 − 0) = 6.0 °C/min. The 200 g line rises 15 °C in 5 min: 3.0 °C/min, half as steep.
- Graph choice: line graph for a continuous independent variable, bar graph for categories, scatterplot for two measured variables, histogram for how one variable is distributed.
- Reading between points is interpolation (usually safe); reading beyond them is extrapolation (risky: water heated past 100 °C boils instead of getting hotter).
- A correlation shows two variables change together. It does not show that one causes the other; a controlled experiment tests that.
Time goes on the x-axis and temperature on the y-axis. A line graph fits; categories such as "enzyme A, B, C" would need a bar graph instead. You choose an even scale (every 10 °C, every 1 min) that fills most of the grid and label each axis with its unit. A key is needed so a reader knows which line is 100 g and which is 200 g. You draw a trend line and find its slope (rise ÷ run, here 6 °C/min) to compare how steeply the groups change.
- x-axis
- The two number lines of a graph. The x-axis (horizontal) normally carries the independent variable and the y-axis (vertical) the dependent variable; each needs a label with units and an even scale.
- line graph
- A graph of points joined in order along a continuous independent variable such as time or temperature, used to show how the dependent variable changes across it.
- bar graph
- A graph of separate bars, one per category (for example, four different enzymes), used when the independent variable is categorical rather than continuous.
- histogram
- A graph of how often values fall into each range (bin) of one measured variable, with touching bars; it shows the shape and spread of a distribution.
- scatterplot
- A graph of unjoined points, each pairing two measured values from one individual or sample, used to look for a relationship between the two variables.
- trend line
- A single straight or smooth line drawn through the overall pattern of the points (a line of best fit), not joining each point.
- correlation
- A pattern in which two variables tend to change together: positive if both rise together, negative if one rises as the other falls.
- causation
- A change in one variable producing a change in another. A correlation alone does not show causation; a controlled experiment is needed to test it.
- choosing a graph type
- Choosing the graph type from the data: line graph for a continuous independent variable, bar graph for categories, scatterplot for two measured variables, histogram for the distribution of one variable.
- graph key
- The key (legend) on a graph that tells which line, symbol or bar color stands for which group or condition.
- interpolation
- Reading a value between measured points (interpolation) or beyond the measured range (extrapolation). Extrapolation is risky because the pattern may not continue.
- slope
- How steep a line is: the change in y divided by the change in x between two points (rise over run), with units of the y-unit per x-unit.
Standard deviation and standard error
The standard deviation, s = √[Σ(x − x̄)² ÷ (n − 1)], is the typical distance of a value from the mean, in the data's units. In a normal distribution about 68% of values lie within 1 SD and about 95% within 2 SD. The standard error, SE = s ÷ √n, measures how precisely the sample mean is known; it shrinks as n grows, while the SD does not.
- Worked example. 58, 60, 62, 64, 66 s. x̄ = 62. Deviations −4, −2, 0, 2, 4; squares 16, 4, 0, 4, 16; sum 40. 40 ÷ (5 − 1) = 10. s = √10 = 3.16 s. SE = 3.16 ÷ √5 = 1.41 s.
- Formula sheet: s = √[Σ(xi − x̄)² ÷ (n − 1)] and SEx̄ = s ÷ √n. The standard deviation has the same units as the data.
- In a normal distribution, about 68% of values lie within ±1 SD of the mean and about 95% within ±2 SD.
- SD answers "how much do individuals vary?"; SE answers "how precisely do I know the mean?". To halve SE, collect four times as many values.
You find each deviation (x − x̄); positive and negative deviations would cancel, so you square them. This gives the average squared distance from the mean (dividing by n − 1 corrects for using the sample's own mean). The result is the standard deviation, s: a typical distance of a value from the mean. The standard error, SE = s ÷ √n, estimates how much the mean itself would vary if you repeated the sample. Larger samples pin down the mean more precisely, while the standard deviation stays about the same.
- standard deviation
- A measure of how far the values in a sample typically sit from their mean: s = √[Σ(x − x̄)² ÷ (n − 1)]. It has the same units as the data, and a larger s means more spread.
- standard error of the mean
- The standard error of the mean, SE = s ÷ √n: an estimate of how much the sample mean would vary if you repeated the whole sample. It shrinks as the sample gets larger.
- normal distribution
- A symmetric, bell-shaped spread of values around the mean. About 68% of values fall within 1 standard deviation of the mean and about 95% within 2.
Rates and percent change
A rate is change ÷ time (the slope); the average rate over an interval can hide a rate that changes within it. Percent change is (final − initial) ÷ initial × 100, with a sign. Convert units with metric prefixes before calculating, remembering that areas and volumes convert by the squared and cubed factors. Report every answer with units.
- Worked example. Depth 1.6 mm at 2 min, 3.2 mm at 8 min, 3.6 mm at 10 min. Rate 0-10 min = 3.6 ÷ 10 = 0.36 mm/min. Rate 8-10 min = (3.6 − 3.2) ÷ 2 = 0.20 mm/min. Percent change from 0.8 to 0.2 mm/min = (0.2 − 0.8) ÷ 0.8 × 100 = −75%.
- Formula sheet: rate = change in y ÷ change in x; percent change = (final − initial) ÷ initial × 100. Always divide by the initial value.
- The difference between two percentages (97.8% and 56.1%) is 41.7 percentage points; the percent change is (56.1 − 97.8) ÷ 97.8 × 100 = −42.6%.
- Prefixes: kilo 10³, centi 10⁻², milli 10⁻³, micro 10⁻⁶, nano 10⁻⁹. Areas convert by the factor squared, volumes by the factor cubed.
Its average rate is the change divided by the time: 3.6 mm ÷ 10 min = 0.36 mm/min. The rate over a short interval differs from the average: 0.8 mm/min in the first 2 minutes, 0.2 mm/min in the last 2. You calculate percent change = (final − initial) ÷ initial × 100; from 0.8 to 0.2 mm/min is −75%. You convert with metric prefixes before calculating: 1 mm = 1,000 µm, so 1 mm³ = 10⁹ µm³. You report the answer with its units (mm/min, %, cm⁻¹) and the sign, and round as the question asks.
- rate of change
- How fast a quantity changes per unit of something else, usually time: change in y ÷ change in x (for example, mm per minute). On a graph it is the slope.
- percent change
- The change in a value as a percentage of its starting value: (final − initial) ÷ initial × 100. Positive for an increase, negative for a decrease.
- metric prefixes
- Prefixes that scale metric units by powers of ten: kilo (10³), centi (10⁻²), milli (10⁻³), micro (10⁻⁶), nano (10⁻⁹). Scientific notation writes numbers as a × 10ⁿ.
- average rate
- The total change divided by the total time over an interval (the slope of the straight line joining its end points). The instantaneous rate is the slope at a single moment, which can differ along a curve.
95% confidence intervals and error bars
A 95% confidence interval, about x̄ ± 2SE, is a range that very probably contains the true mean; drawn on a graph it becomes an error bar. If two groups' 95% CI bars do not overlap, the difference is statistically significant and the null hypothesis is rejected. If they overlap, the data do not show a difference, which is not the same as showing there is none. Larger samples narrow the intervals. Always label what error bars represent.
- Worked example. 15 °C: n = 9, x̄ = 1.8, s = 0.30, SE = 0.10, CI = 1.6-2.0. 25 °C: n = 9, x̄ = 2.5, s = 0.36, SE = 0.12, CI = 2.26-2.74. No overlap: likely a real difference.
- Formula sheet: SE = s ÷ √n; 95% CI ≈ x̄ ± 2SE. A bigger sample gives a smaller SE and a narrower interval.
- Overlap rule (95% CI bars only): no overlap, likely a real difference; overlap, the data do not show one. Overlap never proves the means are equal.
- Always say what your error bars show (SD, SE or 95% CI). The same picture means different things for each.
You can build a range around it that very probably contains the true mean. The interval x̄ ± 2SE is an approximate 95% confidence interval. A reader can see how precisely each mean is known. The difference is unlikely to be chance: it is statistically significant, and you reject the null hypothesis. The data do not show a difference, so you fail to reject the null; more data could still reveal one.
- confidence interval
- A range around a sample mean that very probably contains the true mean. For the exam, a 95% confidence interval is about the mean ± 2 standard errors (x̄ ± 2SE).
- error bar
- A line drawn above and below a plotted mean to show uncertainty. Bars can show SD, SE or a 95% confidence interval, so the graph must say which.
- overlapping error bars
- A rule of thumb for 95% confidence interval error bars: if the bars of two means do not overlap, the difference is likely real; if they overlap, the data do not show a difference (which is not proof that there is none).
- statistically significant
- A difference is statistically significant when it is too large to be explained by chance alone if the null hypothesis were true, so the null hypothesis is rejected.
Water potential calculations
Water potential is Ψ = ΨP + Ψs. Solute potential is Ψs = −iCRT: i is the number of particles per formula unit (1 for sucrose, about 2 for NaCl), C the molar concentration, R = 0.0831 L·bar/(mol·K) and T the temperature in kelvin. In an open container ΨP = 0. Water moves from higher to lower water potential. 10 bars = 1 MPa.
- Worked example. 0.30 M sucrose, 22 °C, open beaker. i = 1, C = 0.30 mol/L, R = 0.0831 L·bar/(mol·K), T = 295 K. Ψs = −(1)(0.30)(0.0831)(295) = −7.35 bars. ΨP = 0, so Ψ = −7.35 bars (−0.74 MPa).
- Same molarity is not same water potential: 0.1 M NaCl (i = 2) has Ψs = −4.95 bars at 25 °C, twice the 0.1 M sucrose value of −2.48 bars.
- A potato-core graph crossing 0% mass change at about 0.30 M sucrose means the cells' Ψ equals that solution's Ψ: about −7.4 bars at 22 °C.
- Three classic slips: using °C instead of kelvin, forgetting i for salts, and dropping the minus sign.
The concentration of dissolved particles is i × C. Solute potential is negative and proportional to particles and temperature: Ψs = −iCRT. T must be in kelvin (°C + 273), and the answer comes out in bars. Water potential is Ψ = ΨP + Ψs, with ΨP = 0 in an open container. Net water moves from the higher (less negative) Ψ to the lower (more negative) Ψ until they are equal.
- solute potential equation
- The formula-sheet equation for solute potential: Ψs = −iCRT, where i is the ionization constant, C the molar concentration, R the pressure constant and T the temperature in kelvin.
- ionization constant
- In Ψs = −iCRT, i is the number of particles each formula unit makes when it dissolves: 1 for sucrose or glucose (they do not ionize), about 2 for NaCl (Na⁺ + Cl⁻), about 3 for CaCl₂.
- molar concentration
- Moles of solute per liter of solution (mol/L, written M). A 0.2 M sucrose solution has 0.2 mol of sucrose in each liter.
- pressure constant
- In Ψs = −iCRT, R = 0.0831 L·bar/(mol·K). It converts concentration and temperature into a pressure in bars.
- kelvin
- The absolute temperature scale used in Ψs = −iCRT: kelvin = °C + 273. Room temperature, 22 °C, is 295 K.
- bar (pressure unit)
- A unit of pressure used for water potential; 1 bar is about atmospheric pressure at sea level. 10 bars = 1 megapascal (MPa), the unit many textbooks use.
- water potential calculation
- Finding water potential with Ψ = ΨP + Ψs: calculate Ψs = −iCRT, add the pressure potential (0 in an open container), and compare values; water moves from higher (less negative) to lower (more negative) Ψ.
Chi-square test
The chi-square test compares observed counts with the counts expected under a null hypothesis: χ² = Σ (o − e)² ÷ e. Degrees of freedom are categories − 1. If χ² is greater than the critical value at p = 0.05, the deviation is unlikely to be chance and the null is rejected; otherwise you fail to reject it. Use counts, not percentages, and never "prove" or "accept" a hypothesis.
- Worked example. o = 29 moist, 11 dry; e = 20, 20. χ² = (29 − 20)² ÷ 20 + (11 − 20)² ÷ 20 = 4.05 + 4.05 = 8.1. df = 1; critical value (p = 0.05) = 3.84. 8.1 > 3.84: reject the null.
- Formula sheet: χ² = Σ (o − e)² ÷ e, with df = number of categories − 1. Use the p = 0.05 column unless told otherwise.
- Bigger χ² means a worse fit to the null. χ² greater than the critical value means p is less than 0.05: reject. Otherwise, fail to reject.
- Always use counts, never percentages or means: chi-square depends on how many individuals were counted.
That prediction gives an expected count, e, for each category. Adding the terms gives χ² = Σ (o − e)² ÷ e: 4.05 + 4.05 = 8.1. You look up the critical value for df = categories − 1 (here 1) at p = 0.05: 3.84. A gap this large would happen less than 5% of the time if the null were true (p < 0.05), so you reject the null hypothesis. …you would fail to reject the null: the counts would fit "no preference" well enough to be chance.
- observed value
- Observed values (o) are the counts actually recorded in each category. Expected values (e) are the counts the null hypothesis predicts for the same total.
- chi-square test
- A test of whether observed counts differ from expected counts by more than chance would explain: χ² = Σ (o − e)² ÷ e. The bigger χ², the worse the fit to the null hypothesis.
- degrees of freedom
- For a chi-square goodness-of-fit test, the number of categories minus 1. It tells you which row of the critical value table to use.
- significance level
- The cutoff probability chosen before the test, usually p = 0.05: the risk you accept of rejecting a null hypothesis that is actually true.
- critical value
- The χ² value from the table for your degrees of freedom and significance level. If the calculated χ² is greater than it, reject the null hypothesis.
- p-value
- The probability of getting a deviation from expected at least as large as the one observed if the null hypothesis were true. A p-value below 0.05 leads to rejecting the null.