Sunday, June 12, 2011

Evaluating Teachers’ Performance

The Michigan State Legislature has decreed that all public school districts will, beginning this fall, annually evaluate all teachers and principals in a way that considers student growth over time as a significant factor, translating that growth into satisfactory or unsatisfactory ratings of the staff members.

On the face of it, the plan sounds reasonable: why not measure educators’ performance to a significant degree by learning outcomes for students? After all, good learning outcomes are arguably the primary purpose of schooling, are they not?

My question would be whether this is the best way to achieve them. There are other issues, such as the reliability of our testing methods for determining student knowledge and skill levels, the difficulty of evaluating teachers who do not teach tested subjects, and the efficiency of trying to improve teaching with threats rather than professional development and support. But today, I want to consider the simple cost of such evaluations and a different way to spend that kind of money to better achieve our goal of student learning.

What will it cost?

To start with, we can throw out the MEAP tests for this purpose. In order to fairly evaluate student growth over time, we need to test them at the beginning of each year to establish baselines, and at the end of each year to measure growth. To give it an educational as well as an evaluation role, it would also have to be given several times during the year, to enable course corrections and interventions if they are not doing as well as expected.

You can get an idea of what this will cost by what the Ann Arbor Public Schools (AAPS) has just committed to spending for norm-referenced Northwest Evaluation Association (NWEA) testing:

  • $62,856 annually for K–2 testing (plus Scarlett MS)
  • $49,491 annually for grades 3–5 testing
  • $51,611 annually (starting next year) for grades 6–8 testing
  • $31,500 for the new server space this testing will require (one-time expenditure)
  • Unknown but significant amount for new computers required for middle school testing
  • An estimated $50,000 in two years to pilot software for another evaluation model

That amounts to about $164,000 per year in just K–8 testing costs, after a significant investment in computers and servers to handle it, followed by another significantly expensive software pilot, presumably leading to a much more expensive broader implementation later. This is not chump change. Can you say, “unfunded mandate”? Note that districts are supposed to come up with the funds for this evaluation system at the exact same time that our per-pupil funding has been cut by $470 and our retirement costs have increased by 18% (including the one-year reprieve in rate increase).

My understanding of the NWEA testing is that it is actually efficient and useful. Because it responds dynamically to student input, the difficulty level adjusts up or down to give a good estimate of just what each individual student knows and can do. And because it is computerized, with nearly instant results, it provides specific information that is useful to teachers.

The next step

But testing — whether for teacher evaluation or to allow timely adaptation by teachers — is only part of what we need. It may diagnose a problem but does not solve it. Suppose a child is not learning or a teacher does not seem to be teaching effectively — then, what? To pursue our goal of student achievement, we need ways to help teachers do their job better: professional development that works. Montgomery County, Maryland, where I grew up, has an innovative way of coaching teachers who need help. Its Peer Assistance and Review program formally mentors both new teachers and veterans who are underperforming, according to their principals’ evaluations.

Intensive help in the form of modeling, planning, coaching, and reviewing instruction is provided for a year by experienced and highly qualified Consulting Teachers. CTs induct new teachers into the school culture, providing practical tips and demonstrating what good teaching looks like. They also provide struggling teachers with intensive support and assistance to improve their practice. CTs are provided the same Observing and Analyzing Teaching courses offered to principals. After a three-year rotation in the CT position, they return to the classroom — with improved leadership and communication skills now available to fellow teachers who are not in such dire need of help.

After these year-long interventions, a PAR panel of teachers and principals reviews the CT data and report and evaluates whether staff are meeting the district’s six standards of effective teaching: commitment to students and their learning; knowledge of their subject and how to teach it; maintaining a positive learning environment; continually assessing student progress and adapting instruction to improve it; commitment to continuous improvement in their practice; and exhibiting a high degree of professionalism. The panel then recommends that teachers be returned to the regular professional growth cycle, be given a second year of support, or be dismissed.

While the point of the system is not simply to find and fire “bad teachers,” teachers who need to improve but do not are, in fact, dismissed. According to a recent New York Times recap of the PAR program, panels have voted to fire 200 teachers in the past 11 years, and 300 more have left rather than go through the PAR process. (Keep in mind that school districts in Maryland are county-wide and therefore very large; MCPS has nearly 150,000 students.) For comparison, in the ten years before PAR, only five teachers were fired.

What is more important, though, is that hundreds of new teachers were effectively mentored through their always-difficult first year, and hundreds more struggling teachers were helped to become the effective and professional instructors that all children deserve. Moreover, the program’s careful design, and equal panel representation of teachers and administrators chosen by their unions, have resulted in real trust and buy-in by all parties. The program is seen not as merely punitive but as a genuine opportunity to improve professional practice. Even before that trust was built, a 2004 report found that most tenured teachers who were put into the program were grateful for it afterward, acknowledging that it made them better teachers.

In other words, the PAR program improves student outcomes by improving teacher practice. Yet it does not meet federal standards for school improvement.

Elevating means above ends

Federal school improvement grants awarded through Race to the Top are intended to accelerate student growth, but the program is very specific about just how this must be demonstrated. Specifically, districts are required to evaluate teacher quality by means of students’ test scores, just as Michigan now requires. Montgomery County PS, already getting results the rest of us would envy, had to turn down the $12M it could have gotten from Maryland’s RttT grant. As its superintendent noted, “We don’t believe the tests are reliable. You don’t want to turn your system into a test factory.”

Well, that horse is already out of the barn here, I’d say. Tests are moving, metaphorically, from “the important thing” to “the only thing.” We seem to have forgotten their point.

Another way

Suppose, instead, we diverted some of the additional money we will now have to spend on testing to a proven professional development model like PAR. A handful of master teachers delegated to coaching new and struggling staff on an ongoing basis could do more than define the outlines of a problem, as testing does. Instead, they could actually solve it. Building better teachers produces better student achievement. Is that so hard to believe?

Wednesday, April 27, 2011

Fooling Ourselves

How our motivations remain hidden even from ourselves

I have been reading David Brooks’s masterpiece, The Social Animal, and now see many things with newly opened eyes. Brooks has synthesized a huge amount of recent research into neuropsychology and chose to communicate it in an extremely readable “story” format. This device is an act of pure genius, it seems to me.

We all know the limitations of “do as I say” as a means of teaching life lessons. Every parent has tried and failed to help children learn something the easy way. Simply put, telling others the lessons we have learned does not work — it does not prevent them from learning things the hard way just as we did.

That does not mean, however, that we cannot pass on our accumulated wisdom to others. To do so effectively, instead of telling them the conclusion, we must share the whole story. Narratives immerse us in the full experience that resulted in learning. A good story lets us feel the same emotions that the participants did, and that emotional context is what made the lasting impression on their memory.

This is why we enjoy and can learn from others via novels, dramas, and even campfire story sharing. We go through a vicarious, foreshortened version of the storyteller’s experience. It lets us empathize deeply with those unlike us or experiences we will never have. The story need not even be a long one for this work.

For example, this morning I was reading a newspaper article about something happening at Arlington National Cemetery containing this phrase: “At one grave was a baby's sonogram.” Did you experience the same visceral reaction as I? I instantly filled in the story: some military man’s widow is pregnant, and this is the only way she can share the news with him. I doubt you have to be a former military spouse, as I am, to get a jolt over the life-altering sacrifices going on around us, largely unremarked, every day. This is the power of story.

The point of the book: we don’t really know ourselves

Brooks cites utterly convincing evidence that the majority of our feelings and reactions remain unconscious. We do things for reasons we do not see or understand, and our conscious, rationalizing mind tries to explain them after the fact.

Complex and sophisticated decision-making goes on all the time below the level of our awareness. We constantly sift through the data from our senses for important clues to guide our behavior. Where do I need to be in a second or two to catch that ball? Is that person’s smile genuine or devious? Does the speed and trajectory of that truck give me time to cross the intersection safely? Is that approaching dog a threat?

How we assess and prioritize all this data depends upon our lifetime of previous experiences. The trained athlete can calculate and predict the ball’s behavior. Our vast store of human interaction data helps us judge the authenticity of a smile. A practiced driver knows how quickly that vehicle will close with his. Our past acquaintanceship with dogs informs our understanding and expectations of them. But all these sophisticated assessments are unconscious. We are impelled to act in certain ways without realizing why.

This explains so much. Why we impulsively do things that we know are not good for us. Why we cannot simply resolve, consciously, to change our ways and stick to that decision. Why some folks are confirmed cynics or persistent conspiracy theorists, certain that no one else can be trusted. Why others are naive and overly trusting. What we have experienced, especially repeatedly, is what we expect to happen.

One interesting implication re conflict of interest

Then I read an op-ed piece in the New York Times by Harvard and Notre Dame business professors on ethical blind spots. Max Bazerman and Ann Tenbrunsel wrote the new book “Blind Spots: Why We Fail to Do What’s Right and What to Do About It.” They note how hidden influences on behavior let us behave unethically without realizing it.

From Bernie Madoff on down, many perpetrators of financial fraud have had a disconcerting tendency to rationalize and excuse the most egregiously wrong behavior. Madoff told The Times that banks and hedge funds were “complicit” in his massive Ponzi scheme, that they “had to know” that something wasn’t right. He was deluding himself that he is not responsible, and he was (correctly) pointing out that his fraud was obvious but ignored by others who benefited from it.

The professors cite research that substantiates this idea. Many experiments reveal that a focus on group goals produces an “ethical fading.” We overlook transgressions and conflicts of interest when it is in our interest to do so. For example, long-term attorneys and auditors, having invested in and developed a relationship with clients, are no longer objective in their advice. They unconsciously tell the clients what they want to hear.

The really interesting data reveals that “sanctions, like fines and penalties, can have the perverse effect of increasing the undesirable behaviors they are designed to discourage.” When people face fines for wrongdoing, they tend to cheat more, because they see the situation as a financial rather than ethical dilemma. When there is no fine, they are more conscious of their ethical responsibilities and behave better.

If fines don’t work, how about transparency? In recent discussions about how to deal with conflicts of interest, I have held the opinion that the important thing is to disclose them: as long as they are not hidden, then others can judge our actions and motivations. But other research “found that disclosure can exacerbate such conflicts by causing people to feel absolved of their duty to be objective.” Yikes.

Think of the implications here. Fining BP will not only not prevent a repeat of the behavior leading to the Gulf oil disaster, it may actually encourage it. [I’m not ignoring the absolute benefit of fines to pay for the damages and clean-up.] Having public officials declare their conflict of interest before a vote may actually set them free to vote with bias. Even, dare I say, financial penalties for not achieving high enough test scores in schools may actually encourage cheating.

Bazerman and Tenbrunsel believe we need safeguards that prevent misbehavior rather than threaten to punish it. In their field, they recommend strict division of responsibilities to minimize ethical conflicts. Auditors should only audit. Credit-rating agencies should not be financially intertwined with those they rate. And I’m leaning toward conflict of interest policies that require more than disclosure. Perhaps office-holders should not be allowed to have some kinds of relationships — or should abstain from voting when they do.

Tuesday, March 29, 2011

The Elephant in the Room

Rising poverty and wealth inequality are threatening the health of our nation and our democracy. Why aren't we all talking about this?

First, let's establish where we are

As Paul Krugman noted earlier this month, “one-sixth of America’s workers … can’t find any job or are stuck with part-time work when they want a full-time job…. Unemployment has become a trap, one that’s very difficult to escape. There are almost five times as many unemployed workers as there are job openings; the average unemployed worker has been jobless for 37 weeks, a post-World War II record.”

We are told that cuts in corporate taxes and in government spending will produce the new jobs we so desperately need, but nowhere has this actually been shown to work as advertised. This is classic supply-side economics but, with so many un- and under-employed, our problem is lack of demand. If more people had more money to spend, businesses would be hiring more workers to produce what they want to buy. Business costs are already moderated by historically low interest rates, and the profit levels and cash on hand of the largest corporations (more than $2 trillion, in one estimate I found) have rebounded to pre-crisis levels. Yet they are not hiring in significant numbers, because demand lags.

What about small businesses, often touted as the real job creators in our economy? Entrepreneurs need working capital to get started. From 2001 to 2003, cuts in capital gains, dividend, and estate taxes amounted to more than $3 trillion, yet there is little evidence the proceeds of these cuts (since extended) have been invested in job creation. The Small Business Association of Michigan recently surveyed its members, asking how they would use the money saved by no longer paying businesses taxes, as Gov. Snyder’s plan proposes. Only 48% said they would add employees. While 51% said they would buy new equipment, that can often lead to an actual loss in jobs, as it did in the auto industry. And a full 50% reported they would “realize the profits” — which creates no new jobs, of course.

For at least a generation now, political leaders have preached the religion of supply-side economics, yet wealth and income disparities have grown exponentially.

How’s that “rising tide” working out?

According to the adage, a rising tide lifts all boats — we all benefit when the rich are allowed to get richer. And richer they have gotten. Last December, Sen. Al Franken quoted the Economic Policy Institute to note that “during the past 20 years, 56 percent of all income growth has gone to the top one percent of households. Even more unbelievable — a third of all income growth went to just the top tenth of one percent. At the same time, middle class families have done decidedly worse. When you adjust for inflation, the median household income declined over the last decade.”

The disparity grows, yet Americans habitually and significantly underestimate the extent of wealth inequality in the U.S. A nationally representative random sample of respondents surveyed by Harvard Business School researchers “vastly underestimated the actual level of wealth inequality in the United States, believing that the wealthiest quintile held about 59% of the wealth when the actual number is closer to 84%.”

This is not what we want. In the same survey all demographic groups exhibited a surprising consensus on the “ideal” distribution of wealth. They all approved of some inequality, but their ideal was far more equal than the current level, and “far more equitable than even their erroneously low estimates of the actual distribution.”

What has this to do with education?

Americans believe strongly in offering all the opportunity to do better in life, and they recognize that educational opportunity is the key to changing one’s economic circumstances for the better. Our rising child poverty rates are directly related to our difficulty in erasing student achievement gaps.

Just as we are in denial about growing wealth disparity, we seem blind to the fact that more children are poor. Census data from 2000 showed more than 23% of children in Wayne County lived in poverty. By 2008, that percentage had risen to 29.3%. By 2009, 59% of school-age children in the county qualified for free or reduced-price lunch. [See KidsCount.org for many such statistics.]

It is not an excuse to note that poor children need much more help to overcome the disadvantages of starting further back and having less support at home. Their preparation and support for learning must be greater than that required by more advantaged children, yet funding inequality in education persists and, by some measures, is getting worse. Children from wealthy families routinely enjoy better schools — with more resources, lower pupil-teacher ratios, better-paid staff, and better-equipped buildings — in addition to their advantages at home. Poor children tend to need much more but to get much less.

The results are documented in our national results on the international PISA achievement tests. National Association for Secondary School Principals researchers disaggregated the 2010 results by income and issued a report entitled “PISA: It’s Poverty Not Stupid.” When comparing apples to apples — other nations and American schools with equally low poverty rates — our students were first in the world.

It is our poor students who perform poorly. The reasons why are no mystery: poorer nutrition (starting before birth), poorer health (especially high levels of asthma), poorer attendance (due to poor health and chaotic households), higher incidence of drug use and violence in their homes, higher rates of homelessness, lower quality child care, fewer books in the home, and so on and so on. All of these handicaps affect not just preparation and and support at home for academic achievement but — more importantly — children’s ambitions, motivation, and sense of what is possible for them.

All of those things can be overcome, but not by magic. Blaming and penalizing teachers for not being able to work miracles without extraordinary resources will not help. (Why would anyone want to work with our neediest students if their pay and job security depends upon their working miracles?) Further cutting already-inadequate funding to schools with high-needs students will not help. Turning to charter schools, whose overall track record with poor children is worse than that of public schools, will not help.

What will help is a serious commitment to provide for needy children what their households cannot, so that they can catch up with the wealthier children who start out so far ahead of them. Such a commitment requires money, time, and effort. Anyone who tells you there is a magic shortcut is selling you a bill of goods.

Tuesday, February 22, 2011

Just why are teachers unionized?

As I write on Feb. 22, it has been interesting to view the revolt going on in the Wisconsin State Capitol, with public employee unions and their supporters protesting Gov. Scott Walker’s non-negotiable demand that they be stripped of collective bargaining rights. The monetary issues involved (increased contributions and copays for pensions and health benefits) have already been agreed to by the unions. But the governor insists he will not compromise on essentially killing the unions, which would have no purpose if they cannot bargain collectively over work conditions.

There are several intriguing aspects to this battle. For one, if the state cannot bear the financial implications of unionized public workers, why were police and firefighter unions exempted from the demand? One cannot help but notice that the exempt unions are the ones whose members have been more reliable supporters of Republican causes and candidates, whereas the targeted unions have more often supported Democrats. Union PAC money is almost the only potential counterweight to the now-unlimited funds available to candidates from corporate interests. Recall for a moment just who caused the Great Recession that undermined our security and nearly bankrupted our schools, cities, and states. In an era when income disparity has set new records evoking the Robber Baron period of our history, it is a useful distraction if the majority of us can be set at each other’s throats over who is being pushed to Third World living standards faster. (Pay no attention to the obscenely rich behind the curtain!) Are all these concerns irrelevant to the dispute? Likely not, but they are not my topic for the day.

I want to write about a particular target of the governor’s demand: teachers’ unions.

Why are teachers unionized anyway? Aren’t they professionals, and don’t we associate unions with protection for blue-collar workers from dangerous working conditions? You have to know something about the history of American public education to understand the difference it made to have teachers’ unions.

Pre-college teaching used to be a largely female profession, because teachers were paid so little their income could only supplement, rather than actually support, a household. Teachers who did try to support families routinely took minimum-wage summer jobs for the extra income. (I know this from the teachers I worked with every summer during my high school and college years as a waitress. We earned 25 cents an hour plus tips. They all had at least one master’s degree.) The best and brightest young women were often attracted to teaching when there were few other opportunities for them. Now, in an era with every work sphere open to them, most women with several degrees will expect a professional wage. In right-to-work states like Arizona, the starting salary for public school teachers of about $26,000 certainly discourages people with other options from entering the profession.

But decent wages and benefits are not the only reasons teachers had to organize and bargain collectively.

Before unionization, teachers were treated much like children: their behavior on and off the job had to be above reproach from the most conservative elements of society. Single women teachers, for example, had to live in chaperoned environments such as approved boarding houses. If they lived on their own, who knows what immoral trouble they might get into! Not so long ago, marriage — a hallmark of adulthood — automatically disqualified a woman from teaching. Even after married women were allowed to remain as teachers, they were routinely excluded once they were obviously pregnant.

Beyond those plainly indefensible restrictions, however, the entire system of public education conspired to keep teachers — both male and female — subservient. They were subject to the whims and prejudices of administrators and school boards. A parental complaint about a reading assignment, a grade, or something said in class could get them summarily fired. There was no respect for their professional expertise, which could be overridden by school boards with no educational credentials or experience. (In some ways, that lack of respect is ascendant once more, now that so many believe our most needy and vulnerable students can best be taught by Teach For America participants with five weeks of training and zero experience.) Authoritarian and paternalistic administrators expected quiet compliance and punished “insubordination.” Those we expected to help our children grow into adults with initiative and independence of thought and action were, themselves, treated like children.

If we truly believe that K–12 education is more than just child care, then we must treat its practitioners like the professionals they are. They must be empowered to continuously improve their practice as research and experience show us better ways. They must have the protection of due process to insulate them from the vagaries of public opinion about what and how to teach. The return of public shaming (as when the Los Angeles Times last summer created its own system of ranking teacher effectiveness based upon test scores and published the names of those they deemed ineffective — leading to at least one teacher suicide) is a throwback to that repressive model from an earlier century.

So, even though I have never been a union member nor lived with one and was raised in a very non-unionized part of the country, I can see clear reasons why teachers, in particular, would want to protect the unions that protect them.

But protection of members is not the only useful function unions serve. Educational lobbyists and union PACs are the only powerful advocates to counter the now-enormous influence of wealthy philanthropists and foundations in the fight over public education reform. Few members of the public seem aware of just how much a few extremely rich folks have changed the priorities and methods of reform. Michael Bloomberg and Bill Gates and Eli Broad have no training or expertise in education, for example, but they have exercised barely checked control over where funding for public education is directed. (A brief example: charter schools in New York City are gifted with much higher per-pupil funding and allowed to restrict the number of special-education and English language–learning students, who are dumped in nearby regular public schools, which are then deemed failures for not producing better results with less funding and more needy students.) Parents and other community members may have concerns over these policies and priorities, but they are not organized to effectively advocate against them. Teachers’ unions are. When they are stripped of that capacity, the debate will be completely one-sided. We will all suffer for it.

Sunday, January 23, 2011

Academics Aren’t Everything

It is repeated like a mantra in K–12 education that our focus must be on “improving student achievement.” I doubt that anyone disagrees with that as our primary goal. But I think, sometimes, that our focus on academics alone may be a mistake.

Why? Because it leads us to concentrate too much on curriculum and teaching and not enough on building better students, future workers, and human beings. We need to pay attention to the nonacademic skills and habits that are just as important to success in school, on the job, and in life.

My perspective on this is informed by nearly 20 years of work for a university and for a nonprofit organization in education outreach — that is, outside groups and individuals trying to help K–12 schools achieve their mission. Over those years, these outreach activities evolved from teacher in-service, push-in classroom activities, drop-in tutoring, and group tours and events for children to almost all one-on-one mentoring. Our focus changed as we realized what was really needed: not just changes in curriculum or teaching methods, but changed students.

Over and over, we saw students from the least supportive backgrounds succeeding through sheer determination, while others with many advantages languished. The difference in outcomes was all about personality and character — and traditional academic support does not teach those vital attributes.

University researchers are now demonstrating that “noncognitive indicators” can predict young people’s success in college and on the job. (See, for example, Michigan State University’s Group for Research and Assessment of Student Potential.) The number one predictor of success is “conscientiousness.” When a teen works hard, perseveres through difficulties, shows up regularly, and completes work as expected, he or she will almost certainly do well.

The critical step is that teens must learn to take the initiative, to realize that no one else can determine where they will go in life. Mentoring is one route to that kind of “eureka!” moment. The programs I have worked with have had as their primary goal this transformation of children into self-motivated adults.

While many teens have accumulated appalling deficits in basic skills, they need much more than subject-matter tutorials. They need help with analyzing their motivational problems, with learning to set study schedules, with devising strategies to get out of the deep holes in which they find themselves. Many have no clue, for example, how to deal with being in over their heads except by denial: ignoring homework, cutting classes, and flunking out. They need to be led through determining what can be salvaged, and negotiating with teachers and counselors over the best step to take next: intensified tutoring, makeup work, alternative assessments, changes in sections or entire classes. They need to become convinced that they are not just victims, that there are things they can do to improve their situation. Moreover, they need to learn that only they can rescue themselves. Mentors can diagnose and work on gaps in basic knowledge, can provide hand-holding and confidence-building, but the students must internalize the fact that no one can pour knowledge into them. They are the only ones who can guarantee their own learning. This realization transforms them in a way that no amount of factual knowledge ever could.

I also know, however, that high-quality, long-term, one-on-one mentoring is very difficult to arrange and to sustain. It is not, however, the only solution to the passiveness that holds so many young people back.

We are now learning ways to structure project-based learning and various kinds of team-based learning to encourage just this kind of social-emotional development. The New Tech High program, for example, which we hope to begin at our high school next fall, offers one such format for developing truly college- and job-ready students. The program site I visited last fall in Indiana offered grades for “work ethic” and “collaboration” as well as for content mastery. One could not receive an “A” without behaving in a dependable and conscientious way. Letting down the team prompted conferences that were very like performance reviews on the job, leading the student through a self-analysis of how he or she had failed and what to do differently. The student-created cultural code also discouraged absenteeism, intimidation, theft, and other bad behavior that will not be tolerated in college or work environments.

The key to the success of this program, I believe, is treating students more like adults (and teachers more like professionals, but that’s another topic). People tend to rise to meet higher expectations, after all. Children and teens can grow enormously through taking on more responsibilities. They gain poise and are empowered by success in completing real-work projects and presenting their results. Nothing builds self-esteem and confidence like accomplishing something you thought was beyond you. Challenge and collaboration, when carefully structured, can allow students to risk failure by aiming higher.

You cannot excel if you never try. Trying and succeeding whet one’s appetite for more. Enthusiastic and self-motivated students can’t be held back.

Saturday, January 1, 2011

Teacher “Accountability” Systems

“Value-added” performance measures, which purport to rate teachers by whether a year of their instruction produces a year of learning gains in their students, are the new high-stakes method of deciding who goes and who stays in K–12 education. In pursuit of competitive grants through the American Recovery and Reinvestment Act of 2009 (ARRA), most states (including Michigan) adopted requirements for the use of “effectiveness data” in determining compensation levels for teachers and principals. This data is also, increasingly, being used to justify firing “ineffective” educators. I will limit myself here, for reasons of space, to discussing such performance evaluation schemes for teachers. I have a several concerns about them.

Are they accurate?

Can we all agree that the point of any teacher evaluation system should be to improve the quality of teaching? If so, then such systems must, first of all, be accurate in their analysis of the existing quality. If they do not, in fact, fairly represent the strengths and weaknesses of teacher performance, then they are worthless for the purpose of improving it.

This — the blatant inaccuracy of the evaluations — is the root reason for teacher anxiety and loathing regarding the “accountability” systems to which they are increasingly subject. While many such schemes advocate “multiple measures” of effectiveness, too many are based exclusively on students’ standardized test scores. The error rates in such data are simply unacceptable for high-stakes decisions. I imagine that teachers would not fear evaluations that mistakenly offer them extra help in improving their professional practice, but they would find it unacceptable to lose their jobs over bad data. Wouldn’t you?

Answer to the above question: Student test scores are imprecise measures — they do not fairly and accurately measure student learning outcomes. The U.S. Department of Education’s Technical Methods Report “Error Rates in Measuring Teacher and School Performance Based on Student Test Score Gains” [Schochet & Chiang, Mathematica Policy Research, July 2010, http://ies.ed.gov/ncee/pubs/20104004/] concludes that typical value-added performance evaluations of teachers are unacceptably imprecise. Specifically, “in a typical performance measurement system, 1 in 4 teachers who are truly average in performance will be erroneously identified for special treatment, and 1 in 4 teachers who differ from average performance by 3 to 4 months of student learning will be overlooked.” And that is assuming the use of three years’ worth of data; the reliability is approximately half as good for only one year of data. These error rates, they note, are greatly understated by certain assumptions, such as that students are randomly assigned to schools and to teachers. There are no controls, therefore, for differences in resources among schools or for differences among student cohorts.

The authors reference “findings from the literature and new analyses that more than 90 percent of the variation in student gain scores is due to the variation in student-level factors that are not under control of the teacher. Thus, multiple years of performance data are required to reliably detect a teacher’s true long-run performance signal from the student-level noise.”

Moreover, the paper’s authors note that error rates they analyzed are only one factor (they list six others) that must be taken into account in designing and using appropriately any value-added estimators of performance. Not attending properly to such features in design and application makes the use of such schemes for high-stakes decisions both ineffective in achieving their stated purpose and manifestly unjust.

Do they improve teaching and learning?

The inaccuracy and unfairness of judging teacher performance almost entirely by student standardized test scores is only one problem with value-added performance evaluations. If we are truly interested in improving educational outcomes — and that is the point, isn’t it? — we must also consider the collateral damage they do.

• They undermine collaboration, which effective schools must encourage. Scoring teachers in ways that make them compete against one another inhibits or even discourages the sharing of lore and techniques that can make everyone better teachers. I recall the havoc created under Jacques Nasser at Ford several years ago when engineers were force-ranked in their evaluations (my late husband was coerced into doing some of the ranking then). This zero-sum game, in which one person winning meant another losing, penalized working together and encouraged sabotage of colleagues. It has taken many years to recover from the damage done. Is this the aim of folks who want schools to be “run more like a business”?

• They distort educational decision-making. Which classes students are placed in (should they be challenged or slotted where they can safely deliver higher test scores?); which students are held back or eased out of a school (dumped by charters or shunted into “alternative” schools); what is emphasized in classes (only what is tested); what level of thinking is encouraged (rote memory versus analysis or independent thought) — practices are encouraged that explicitly undermine good education.

• They are contradictory and inconsistent. On the one hand, today’s reformer bloc assumes that anyone can teach — or can administer a school or a district, for that matter. They find solutions in recruiting and placing “outsiders” through Teach for America or via the Broad Superintendents Academy, for example, with brief training and no experience. Yet the punitive teacher evaluation systems are predicated on the assumption that firing “bad” teachers and principals will improve schooling. There is no sense that these people, who have committed time, money, and working years to their professions, can or should be helped, instead, to improve their practice.

I recognize that public education is inherently political, in that we ultimately do what the people want. That does not mean, however, that those of us on the inside cannot try to change the tenor and direction of the public conversation. I deeply believe that the punitive tone of today’s education reform talk undermines the putative aims of reform. Blaming allegedly incompetent educators for all that is wrong in our society may get the rest of us off the hook (and, I add cynically, may serve the unstated interests of certain politicians and of certain corporations that benefit enormously from public spending on testing, textbooks, software, distance learning, et cetera ad nauseum), but scapegoating will not fix those problems. Never has and never will.

What would work?

I cannot stop at stating the problem but must at least make reference to my own preferred solutions. Teacher performance schemes are not all the rage just because we, collectively, want to blame teachers for things not under their control. All of us and all of our children have experienced at least one truly terrible teacher in our lives, and all of us would like to spare others that experience. So, how should teachers be evaluated?

Classroom observation is a time-honored way of seeing how teachers actually teach, even if it is rarely done often enough or with enough specific, useful feedback to make a difference. Teacher Larry Ferlazzo writes that effective observers “know our school, our students and me — and [have] judgment and skills I … respect. I know they are genuinely concerned about my professional development. They understand that helping me improve my skills is the best thing they can do to help our students…. These purposeful visits have produced detailed and helpful feedback that has made me an even better educator.” The Mathematica report notes that “value-added measures and principals’ assessments of teachers, in combination, are more strongly predictive of subsequent teacher effectiveness than each type of measure alone” [emphasis added].

Multiple data points regarding student assessment are required for valid measurements of student learning. A single set of pre- and post-tests annually is insufficient. Teachers working together to create common curricula and assessments for the same classes can not only assure that all students get the same exposure to a subject, but also help one another jointly improve their classroom practice to achieve similarly high levels of learning. This kind of professional collaboration actually does produce the result we say we want from performance evaluation systems: teachers helping one another to consistently become better at their craft, so that students can perform better not just on tests but in life.

For the point of the evaluation is not just to find out how well or how poorly teachers are performing. It is — or at least it should be — to make them better. If we do not believe that teachers, too, can learn, then we do not believe in education at all.

Note: the statistical mathematics used in the Mathematica study cited goes beyond any “expertise” I developed in a single statistical methods course 40 years ago, so my analysis is subject to error. The conclusions that are not direct quotations may be inadvertently misrepresentative.

Publication URL corrected 8 Jan 11.

May 2011 addition: you have got to see this animation summarizing Dan Pink’s Drive work on motivation and incentives!

Monday, November 1, 2010

So, how DID Finland do it?

I’ve written before about how we should emulate the way Finland pulled itself up, from a stultified, centrally controlled, Soviet-Era educational system to one in which all children perform spectacularly on valid international tests of student achievement. Here, I’ll share what they did and how.

Their success is real

First, let’s establish that their success is real. The Programme for International Student Assessment (PISA) tests are sponsored by the 30-member Organization for Economic Cooperation and Development (OECD), although even more non-OECD nations now participate. These tests in reading, mathematics, and science were administered to 15-year-olds in 41–65 countries in 2000, 2003, 2006, and 2009 (results expected in December 2010). Each testing cycle focuses on one of the subject areas, with minor assessments of the others. The international average given is for the OECD (that is, developed) nations.

PISA tests are the gold standard of achievement testing. In order to do well on them, students cannot simply recall factual knowledge; they must extend what they know into unfamiliar settings to solve problems. Students who do well on them have also done well in life — the ultimate test of the success of schooling. In the 2006 cycle, Finland’s students were the top performers in math and science and number two in reading. (For comparison, U.S. students ranked 25th of 30 for math, 21st of 30 for science, and last in reading.) In the 2003 cycle, Finland was at the top in all three categories.

Student success is consistent across all schools

Differences in student performance within schools are generally taken as reflecting natural variation in inborn talent. Variation between schools, however, is an indicator of social inequality — in most places, there is some segregation by socio-economic class among schools, and the wealthier students’ schools have more resources. PISA offers sophisticated statistical analysis of performance controlling for class. Analysis of data shows that, unlike ours, many nations’ students achieve not only high average scores but also scores that do not vary much by socio-economic status: they have realized the ideal of uniformly high student achievement that No Child Left Behind (NCLB) set for us. Again, Finland is tops when it comes to equality of opportunity and performance. And, please note that Finland is no longer homogeneous: recent immigrants, mostly from poorer countries, speak more than 60 languages. In some urban schools, half of the students are from immigrant families whose native tongue is not Finnish.

Here are some examples of results, comparing Finland and the U.S. In the 2006 Science test, mean performance was 563 for Finland (the top performer that year) and 489 for the U.S. (and 500 for the average of OECD nations, on a 1000-point scale). An indicator of unequal opportunity leading to unequal results is the “between-school variance explained by the index of economic, social, and cultural status of students and schools.” This variance was 19% of the total variance in tested countries for the USA but only 1% for Finland. In the 2003 Mathematics test, mean performance was 544 for Finland and 483 for the USA (and 500 for the OECD average). The between-school variance explained by these factors was, again, 19% for the USA and less than 1% for Finland. Even without any controlling for these variables, performance varies by less than 5% overall among Finnish schools. Nearly ALL their students do quite well.

This success does not cost a fortune

I am always wary of examples of success in schools that spend so much more than average as to be irrelevant in the real world of shrinking funding for schools and other civic priorities. For example, the SEED school in Washington, DC, that was put on a pedestal in the recent documentary “Waiting for Superman,” is a residential school that spends $35,000 per student per year — an unattainable ideal for the vast majority of students.

But Finland has transformed its public schools to offer equal opportunity and to achieve consistently excellent results without spending a fortune. OECD figures show that Finland’s total expenditures on educational institutions, expressed as a percentage of Gross National Product, have actually declined over recent decades. Pasi Sahlberg, a Senior Education Specialist at the World Bank and an Adjunct Professor at the University of Helsinki, produced an analysis that shows this, in “Education policies for raising student learning: the Finnish approach.” He found only a weak correlation between student performance and “cumulative expenditures per student from age 6–15.” This figure for Finland was at the OECD average, whereas U.S. spending was about 20% above that average. Note, of course, that Finland’s comprehensive social welfare system provides much more support to less wealthy families than does the American system.

AN ASIDE: The dataset and tools for manipulating it at http://pisacountry.acer.edu.au/index.php are truly amazing. I encourage you to look into them yourself. When news media do reports on PISA results, as they surely will in December, their summaries and assertions may not be valid reflections of the data, but you can judge for yourself. PISA provides a Data Analysis Manual. You can even “Take the Test” of sample questions yourself to see how much more is demanded of students today! Or, use the Dept. of Education’s International Data Explorer interface at http://nces.ed.gov/surveys/pisa/idepisa/. If you feel under-prepared to make statistical judgments, a report on US performance in 2006 from the National Center for Education Statistics can be found at http://nces.ed.gov/pubsearch/pubsinfo.asp?pubid=2008016.

* * *

To summarize, results on this well-respected test series demonstrate not only that Finland’s students are top performers, but also that their performance varies much less by economic, cultural, and social status than in our country. And they do not “throw money” at education to get these results.

So, what do they do differently that we could learn from?

Sahlberg, cited above, notes that Finland has not adopted the “market-oriented reform strategies” that have overtaken much of the rest of the developed world, including the U.S.: “Consequential accountability accompanied by high-stakes testing and externally determined learning standards has not been part of Finnish education policies.” Instead, cultivating leadership, emphasizing teaching and learning, encouraging creativity, ensuring equity, enabling and trusting teacher professionalism, and “intelligent accountability” have wrought this “miracle.” It was accomplished beginning in the 1990s, when Finland was undergoing “a severe economic decline characterized by a major banking crisis.” Sound familiar? Yet, ten years into the process, Finland began to be ranked by the World Economic Forum as one of the world’s most competitive economies, a distinction it has maintained since then. It is also ranked as one of the least corrupt nations — something we may have difficulty emulating.

We do not have the culture of trust, the respect for public institutions, and the shared values of honesty and equity that Finns enjoy, and that enabled their progress in transforming their educational culture. But, clearly, this values-based approach has worked for them — and reform efforts based on competition and coercion have just as clearly failed here, as elsewhere. Student performance on valid tests such as the PISA ones has not improved, despite the nationwide push that began with NCLB and continues under the Obama administration education policies. Collateral damage from these policies continues to increase, as measured by student drop-out rates, cheating on high-stakes tests, rock-bottom teacher morale and steady defections from the profession, elimination of flexibility and creativity in teaching methods, and a narrowing of curriculum down to the most basic of “core” subjects. What we are doing is not working. What do we have to lose by trying an alternative path?

The first step, it seems to me, is a national conversation about shared beliefs and values. No Mission Statement can make any sense until we establish that context first. Do we truly believe that all children can learn, given enough time and the right support? Do we truly believe that all children should have the same opportunity to succeed, which opportunity must be embodied in the uniformly excellent teachers, facilities, programs, and resources available to them? Do we truly believe that a well-trained and experienced teacher actually has some professional expertise worthy of our respect? Do we truly believe that a well-educated and superbly functioning adult needs more than the “three R’s” to get there? Do we even believe that our society will be better off if all of our adults are allowed and enabled to meet those standards? These are the kinds of beliefs and ideals that underlay the Finnish success.

The policies that flow from such consensus on values would be very different from what we have now. It might seem impossible to reach national consensus on anything, at the moment. But it most assuredly will never happen unless we begin to have these conversations. And we could make a start on the process at the state level, meanwhile. Suppose we subsidized higher education for teachers, to include the master’s level required for all Finnish teachers, in exchange for their teaching in Michigan’s schools for a certain period of time? We could make the payback period shorter for work in our needier schools and communities. This subsidy should be offered on a competitive basis, to attract our best and brightest to this important profession. And suppose we got back to reducing the inequity in per-child funding in our public schools from one district to the next? Suppose our state mandates for days or hours in the school year made allowance for time for teachers to collaborate and to develop their skills on an ongoing basis? Suppose we lobby our political representatives to allow us to divert some of the absolute fortune we spend on mandated testing of nearly all students in most subjects every year to targeted testing with results that are available to teachers in time to be used to actually guide instruction?

Are these ideas really so crazy? Does it make more sense to keep doing what is not working?