Showing posts with label calculus reform. Show all posts
Showing posts with label calculus reform. Show all posts

The University of Illinois at Chicago goes on a fishing trip


Title: Subsequent-Grades Assessment of Traditional and Reform Calculus
Author(s): Judith Lee Baxter, Dibyen Majumdar, Stephen D. Smith
Publication type: article
Online: link

Until 1994-1995 the University of Illinois at Chicago taught first year calculus in the `traditional' manner. In the year 1994-1995 they e.g. used Stewart's book. From the year 1995-1996 on they have switched to `reform' calculus for this first year calculus sequence, using the book by Hughes-Hallett et.al..

In this study the authors examine the results of this change. The grades of the 1994-1995 calculus class are compared to those of the 1995-1996 calculus class. Even though the ACT and mathematics placement scores of both classes are very similar, the reform class got significantly higher grades in calculus. The authors also study grades obtained in other classes by these students (and consider this to be a better measure of success than the calculus grades). The authors do not make any firm explicit conclusions, but the overall impression this reader got while reading the article is that they think that their study shows that the reform approach works better. This is confirmed by the following. The authors state

We just wanted to assess the progress of the changeover, to decide whether to continue with the new method.

Since the University of Illinois at Chicago indeed continued with the `new method' we must assume that they considered that the assessment showed that the progress was satisfactory. As I will point out it is at least debatable whether this is indeed the case.

In the article it is said:

We measured many other science and engineering courses available to us.

It is not mentioned how many, but it seems that this might be as high as a few dozen. The authors mention that for six courses the difference between the two approaches is statistically significant at the 5% level (4 in favor of reform and 2 in favor of traditional). Now you should remember what statistically significant means. Statistically significant at the 5% level means that such a difference or a bigger difference happens by chance only in 5% of the cases. If you consider a few dozen cases (which this study seems to do), then you will find a couple of cases that are statistically significant at the 5% level by chance alone by definition of what `statistically significant' means! Statisticians have a word for `research' like this: it's called a fishing trip. There are statistical methods to deal with this issue, the simplest one being the Bonferroni method which amounts to dividing the 5% by the number of cases that you consider.

One of the subsequent courses is studies in-depth: Physics I. The authors only study the students who take this course immediately after the first calculus course (more on this later). The authors claim that their analysis shows that the reform students significantly outperformed the traditional students. This is however very questionable. The authors earlier state about the calculus sequence

Another obviously important variable is the individual instructor. In our 2-year pool, by necessity only about 11 faculty and a similar number of TA’s worked with each sub-pool, so it would be extremely difficult to completely eliminate this factor. As earlier noted under "Background", the department intended a neutral selection of high-quality instructors, sequence-wide. Ideally also the more substantial size of our student pool will provide some compensation for the influence of individual teaching style.

They completely forget about this when they study the results in Physics I. I do not have the data for Spring 1995 and Spring 1996 (the semesters of the study), but in Spring 2007 there were two sections of Physics I, with the same instructor! So it is very plausible that the difference in scores between the traditional and the reform students in Physics I is due to the teaching/assessment differences in this course between Spring 1995 and Spring 1996 and has nothing to do with reform versus traditional in the calculus sequence. As all students know: some teachers are good, some are bad; some teachers give on average high grades, some on average low grades. The teacher factor cannot simply be ignored in the grade-analysis of Physics I.

There is another curious thing about the Physics I analysis. As mentioned the main analysis is only done on the students who take Physics I directly after Calculus I. There is a difference of .564 in mean and .366 in LSMean (table 2 and after some calculation for the mean also by table 4). This exact same number of .366 is also given in table 6 but here supposedly it refers to all students who took Physics I up to Spring 1997. At least that's what I concluded from the sentence `we consider courses through spring semester of 1997'. Are the numbers in table 6 for the other courses also only for those students that took that course at the earliest possible moment? That's never said. Or is this a copy and paste error?

Summarizing, this paper does not show that either reform or traditionally taught calculus is better. The main faults are

  • The authors test multiple hypotheses without adjusting the significance level for each individual hypothesis. This is bound to lead to false positives.
  • The grade of each student is treated as if it is independent of the grades of the other students. This is not true because of the instructor effect.
  • Something fishy might be going on with the data that the authors decide to include and the data they decide to exclude.

Calculus reform: the case of Roger Williams University

Title: Does calculus reform work?
Author(s): Joel Silverberg
Publication type: article
Online: link

The author describes the results of an experiment at Roger Williams University where some sections of the first course in the calculus sequence were taught using a 'reform' approach and the other sections were taught using what the author calls a 'traditional' approach.

The author mentions the textbook used in the reform section (Dubinsky-Schwingendorf) and describes the teaching in these sections (which it seems was all done by himself), but does not mention the textbook(s) used in the so-called traditional sections or what kind of teaching actually went on in these sections. He does mention that there was a weekly computer lab associated with the so-called traditional sections. This alone would qualify them as reform and not as traditional in many eyes. So we might actually be look at reform versus reform here.

The test used to assess the experiment was a common final. The following is what the author has to say about this.

The instructors of the traditional sections prepared a final examination to be taken by all calculus sections. The exam was designed to cover the skills emphasized in the traditional sections rather that the types of problems emphasized in the reform sections. In an attempt to minimize any variance in the ways individual instructors graded their examinations, each faculty member involved in teaching calculus graded certain questions for all students in all sections.

The experiment lasted three semesters, each semester starting with a new population of students. The following information on the grades on the final exams is given.







Final Exam Grades/ Traditional Sections
Percentile Range:0-2425-4950-7475-100
semester 1CCBA
semester 2FCC+A-
semester 3DCBB








Final Exam Grades /Reform Sections
Percentile Range:0-2425-4950-7475-100
semester 1FDCB-
semester 2FC-CB
semester 3CBBB+

We are not given the number of students or the precise number of sections involved.


The above indicates that the performance of the students in the reform sections was worse in the first two semesters of the experiment and better in the third semester of the experiment. The author credits changes that he made in implementation for this. There is of course another possible explanation: it could be that the students who took the third semester version of the reform section were better. It seems that the students were not randomly assigned to sections, so this is something to take into serious consideration. From the article it can be deduced that Roger Williams University, like most universities, knows the SAT scores of its students and has a mathematics placement examination. This data can be used to correct for initial differences between the students in the reform and the traditional sections. This is however not done. Due to this basically no conclusions can be drawn from this study.

As we have seen, this study has several significant flaws. Summarized:

  • It isn't clear whether the sections that are labeled traditional are really traditional.
  • The reform sections seem to all have been taught by the same professor. So it could be that it is just this professor's teaching ability versus that of his colleagues that is measured.
  • The number of students involved is not indicated.
  • No effort is made to adjust for initial student differences. Since the assignment of students wasn't random, this is detrimental to the study.

So the question 'does calculus reform work?' can unfortunately not be credibly answered by this study.

Calculus reform in Minnesota



Title: Redesigning the calculus sequence at a research university:
issues, implementation, and objectives
Author(s): Harvey B. Keynes and Andrea M. Olson
Publication type: journal article
Online: link (full-text for subscribers only)


Authors' abstract. The paper discusses the progress and challenges of a new reformed calculus sequence for science, engineering, and mathematics students developed by the Institute of Technology Centre for Educational Programs and School of Mathematics, University of Minnesota. The main objective of the Initiative is to enable undergraduates to better learn calculus and the critical thinking skills necessary to apply it in a variety of science and engineering problems. Changes in content and pedagogy are emphasized, including instructional teamwork and student-centred learning, involving students working cooperatively in small groups and exploring mathematical ideas using appropriate technologies. Achievement and retention of Initiative students are compared with a control group from the standard calculus sequence. Student attitudes about the usefulness of the Initiative's curriculum, pedagogy, and its influence on learning are discussed. Future implications including new uses of distributed learning are also addressed.

The University of Minnesota performed an, in principle interesting, experiment on calculus teaching. The usual way calculus is taught at a research university is as follows: the students have three hours of lectures a week by a professor and one hour of discussion led by a teaching assistant (usually a graduate student). One can of course investigate whether this is the optimal mix or not. In the experiment the University of Minnesota traded one hour of lectures a week for two hours of discussion a week (and renamed the discussion 'workshop'). This is from the student point of view. From the university point of view the situation is somewhat different since there are multiple workshop sessions for one lecture session. The typical situation is depicted in the following tables.










traditional#students#hourstotal#
lectures10033
workshops2514











experimental#students#hourstotal#
lectures10022
workshops25312

So from the university point of view 1 hour of lectures is replaced by 8 hours of workshops. Even though teaching assistants are cheaper than professors, this will mean an increase in cost. The authors of the study indeed write

Even after implementing all reasonable economies, there is an incremental cost difference of 20-25% over the standard calculus sequence

This is probably an underestimate since in the actual experiment the professor was also supposed to help during some of the workshops and the teaching assistants were supposed to be present during the lectures. The indicated cost increase does seem to be somewhat realistic if this team-teaching is abolished and only the trading of lectures for workshops is considered. Apart from the increased cost there is another problem that the authors mention: staffing. The experimental condition requires three times as many teaching assistants. One of the ways in which this was addressed in the experiment was by employing high-school teachers and undergraduates as teaching assistants. This of course raises all kinds of issues.

Lets look at the results: does trading an hour of lectures for two hours of workshops actually lead to better results? The authors claim that it does, but I'm not convinced. Comparisons were made between the experimental classes and the traditional classes. Students were not randomly assigned to one of the two conditions. It is justifiable to not do this, but then one should be very careful in making comparisons. The authors look at the average calculus grade point average: this is 3.27 for the experimental condition and 2.85 for the traditional condition. They also looked at how many students took a second year of calculus: this was 77% in the experimental condition and 56% in the traditional condition. The authors also gave partially the same questions on the final exam for both conditions. In the experimental condition 76% of the responses to these seven common questions was correct whereas in the traditional condition only 60% was correct. So this is all clearly in favor of the (more expensive) experimental condition. Now comes something strange. The authors compare the grade point average of the students in all their upper division courses. In the experimental condition 43% had a GPA of 3.5 or higher and only 23% had a GPA less than 3.0. For the traditional condition these percentages are 15% and 58%, respectively. From this the authors conclude that the experimental condition provides students with strong mathematical skills necessary for success in future courses. The more obvious interpretation of this difference is that the students in the experimental condition were just smarter. Remember that students were not randomly assigned! And it becomes even stranger. Like all mathematics departments at research universities the University of Minnesota has a mathematics placement test that is administered to all incoming students. This can serve as a pre-test to determine whether the two groups are comparable and to statistically adjust scores if they aren't. This is however not done. The only reason that I can think of that this is not done is that the placement scores are similar to the upper division GPA scores (which shows that this difference is indeed due to the fact that the students in the experimental condition are smarter) and that if one adjusts for this, then the experimental condition turns out not to outperform the traditional condition. At this point it is good to say that the authors of the article are involved in the experiment and therefore have much to loose if the experiment is deemed a failure.

We can do some ballpark statistics on the above information on GPAs. We assume that students with a GPA of more than 3.5 on average have a GPA of 3.75, those with a GPA in between 3.0 and 3.5 have on average a GPA of 3.25 and those with a GPA of less than 3.0 have on average a GPA of 2.5. Then the average GPA of students in the experimental condition is 3.29 and that of the students in the traditional condition is 2.89. Both of these figures are extremely close to the average calculus GPA for these conditions. So the calculus grades seem to fit in perfectly with the grades in all other courses.

Based on the information on the website of the University of Minnesota something can be said about the aftermath of this experiment. Both the (now no longer) experimental and the traditional condition still exist at the University of Minnesota. The traditional sequence seems to attract twice as many students.