Wednesday, March 03, 2010

Personnel Psychology, Spring 2010: SJTs, affect, and job offer timing


The Spring 2010 (v.63, #1) issue of Personnel Psychology is out. Let's look at the highlights:

First out of the gate, a great meta-analysis for anyone interested in situational judgment tests (SJTs; and who isn't?). Christian, et al. looked at 84 studies and found some pretty interesting things:

1) SJTs reported in the literature have been used to measure a variety of things, including leadership skills (37%), some type of composite (33%), interpersonal skills (12.5%), personality tendencies (9.6%), teamwork skills (4.4%) and job knowledge/skills (3%).

2) Criterion-related validity depends--as you might expect--on the match between predictor and performance measure. Conscientiousness measures, for example, predicted task performance much better than managerial performance (rho=.39 and .06 respectively). The highest correlations (albeit based on relatively small samples) were for teamwork skills and personality composites predicting task performance (.50 and .45 respectively).

3) Video-based SJTs tended to have stronger criterion-related validity values compared to paper-based measures. This was particularly true when measuring interpersonal skills (.47 compared to .27).

Second, a small but interesting study by Johnson, et al. on the relationship between trait affect (i.e., being generally disposed to feeling positive or negative emotions) and job performance. Results from 120 matched employee-supervisor pairs from a variety of jobs using both explicit (survey) and implicit (word fragment completion) measures of affect found substantial correlations, particularly between positive affect and performance (in the .50 range), and particularly when using implicit measures.

Something to add to a selection battery, perhaps? Could be perceived negatively by applicants, however, and I can see some questions being raised about the link to medical issues. But the same types of concerns were originally leveled at personality tests and were mitigated by creating measures specifically tied to work behavior. Definitely an area for more research.

Third, check out this study by Becker, et al. on the impact that job offer timing has on acceptance, performance and turnover. The authors found (using data from a Fortune 500 engineering technology company) that for both student and experienced samples, faster offers were associated with higher acceptance rates. Specifically, for experienced candidates, the difference between 2 weeks and 3 weeks taken to make the offer was substantial, whereas for the students 3 weeks versus 4 weeks was important. But, no differences were found in terms of either performance ratings or turnover among employees hired through different offer speeds.

Implication? The study suggests that offer time does impact the likelihood that the offer will be accepted, but viewed broadly this may not have long-term impacts in terms of how employees do on the job. Maybe in cases of good candidate-employer fit, candidates are willing to wait.

Last but not least are the book reviews. Two books are particularly relevant for us, The Structured Interview (Pettersen & Durivage) and Outliers (Gladwell). The first is received very positively and sounds like a great source for anyone wanting more details about the support for and use of structured interviews. The latter is "well worth [a] few evenings" but requires you to overlook the lack of evidence and convenient inferences.

Final notes: those of you interested in multisource performance ratings should check out Hoffman, et al.'s article, which reinforces the impact of having raters from different levels. Chuang and Liao's article also includes a useful measure of a high-performance work system.

Tuesday, March 02, 2010

IPAC Conference + Innovations Award


What are you doing July 18-21? I assume if you enjoy good weather, good company, and--most importantly--great information on state-of-the-art selection practices, you'll be joining me at IPAC's annual conference in beautiful Newport Beach, California.

If you're not only going but have something to present, by all means respond to the call for proposals. It could be a workshop, panel discussion, symposia--pretty much any format you can think of. Don't wait too long, the deadline is this Friday, March 5.

And speaking of the conference, IPAC has announced that nominations for the Innovations in Assessment Award are being accepted from 5/17-6/18. The winner receives not only formal recognition (and bragging rights), but a free pass to the conference.

IPAC's a great group, full of people that are extremely knowledgeable and passionate about using the best selection practices to get organizations the talent they need. Plus, it's the only international (or national) organization I know of devoted exclusively to the topic.

Hope to see you there!

Friday, February 26, 2010

Book review: Strategy-Driven Talent Management


A thought-provoking collection of essays and ideas; but it won't solve all our problems.

The value of a book lies as much with the reader as it does with the content. A book about advanced programming does little good to the person who has problems turning a computer on. A collection of cooking recipes is largely useless to someone who exclusively uses a microwave.

The same is true about business and HR books. Depending upon who you are and where you're at in life, some books may help you, some may be beyond your reach. Such is the case with SIOP's latest entry into its Professional Practice series, Strategy-Driven Talent Management: A Leadership Imperative, edited by Bob Silzer and Ben Dowell.

The book (tome, actually, at nearly 900 pages) is full of thought-provoking pieces from a variety of authors, including some familiar faces such as John Boudreau and Allan Church. There are academics present, but the majority of authors are practitioners in private sector organizations, such as Aon, Ingersoll Rand, HP, Sara Lee, Merck, and Bristol-Myers Squibb.

The book is roughly broken up by major topic area, although the distinction can be hard to maintain. There are chapters on recruitment, executive onboarding, engagement, measurement, and global issues. There's even a 40-page annotated bibliography. But the editors do an admirable job of keeping the topics all related to the broad field of talent management (TM), which they define as, "an integrated set of processes, programs, and cultural norms in an organization designed and implemented to attract, develop, deploy, and retain talent to achieve strategic objectives and meet future business needs" (p. 18).

The book is described as a "comprehensive [collection of] state-of-the-art ideas, best practices, and guidance." It shines on the first two but failed me on the last, although not for lack of trying. The problem is the book is so long, full of so many ideas and case studies, that it's very easy to get lost and not come away with any clear guidance based on the consensus of authors. To some extent this is endemic in any collection of works by separate authors, but it's clearly a collection of "what you might do" rather than a solid prescription for "how to", although some authors do a better job than others.

Another problem is that many authors seem to presume that current TM practices are sub-optimized because they aren't linked to business strategy and results. This may be true if the process is based on non-validated assumptions, but as long as there is a link between job success and specific practices, we're already there. We just haven't made a particularly good link between job success and organizational success, which may explain the attraction to concepts like competencies (mentioned many times in the book).

But my main problem, and this goes for the book as well as the field, is that it treats the concept of talent management as a logical process to be managed. Somewhere in the transition from HR to TM, we lost the H--human. Talent management (and HR) is messy because it involves people. It's political. It changes every day. And you're dealing with emotions, not lines of code. The real challenge--which is discussed but to my mind not driven home--is how to get the talent mindset into the organizational DNA.

There is value to thinking broadly and philosophically about the topic. It helps us plan. But what people really need are concrete suggestions for establishing a self-sustaining high-performance system. In order to do this, we must address the fundamentals (the basic needs of Maslow's hierarchy, if you will), such as:

- HR must learn "the business" and stay close to their customers
- Supervisors must be selected and trained with their talent management role at the forefront
- Success in HR must be defined and measured. It must be communicated, understood, and valued
- Sustained attention to HR success and significant resources must be expended by both HR and line managers

The book does a passable job of presenting these, but you may have to dig for them. The bigger problem is that there seems to be an assumption that what keeps organizations from having a top-notch TM system is a lack of understanding, either of the organizational strategy or best practices in TM, rather than the very real daily troubles that organizations experience, such as:

- Supervisors that hire people they know/like rather than the most qualified person
- People placed into HR with little or no background, interest, or passion for it
- Insufficient resources devoted to TM/HR
- HR managers who are just that--managers--rather than real HR leaders (Avedon and Scholes present a great assessment in Chapter 2 that helps separate these)

Until organizations have these types of "minor"--but real--flaws ironed out, all the charts and good intentions in the world will have very little impact.

Finally, I was also disappointed that there wasn't more in here about evidence-based TM and HR (which may say more about the field than the authors/editors, who acknowledge this lack in Chapter 22). The field desperately needs more research to tie the hard science of assessment with the more anecdotal/consultant practices such as recruiting, retention, and performance management. This will require significantly more research using methods beyond surveys in order to show what works and what doesn't. There are some ties to good research in here, but the hole is significant.


To summarize, the book contains a lot to like, particularly for individuals already schooled in this area looking to optimize their shop, or for graduate students seeking to understand the big picture. But for most HR practitioners (and, I would expect, executives), this book is akin to a collection of recipes for advanced Italian cooking--fabulous for those used to making their own pasta, but beyond the reach of those struggling to make their own sauce.

Thursday, February 25, 2010

Webinar on 21st Century Assessment


Went to a pretty darn good webinar yesterday put on by HCI and featuring Ken Lahti (PreVisor) and Charles Handler (Rocket-Hire). The topic was 21st century assessment.

Some of the topics covered included:

- increased functionality and usability of testing platforms

- increased sophistication of security methods

- off-the-shelf tests and "I/O psychologists in a box"

- integrating assessment with your overall talent strategy

And my two favorites:

- advanced simulations (such as those using video game technology)

- candidate data that follows them

The webinar is going to be re-broadcast several times today and tomorrow, if you have a chance check it out. You can also see a copy of the slides for free if you're an HCI member (which is free).

Free, short, and full of information--that's my kind of training.

Saturday, February 20, 2010

Recruitment v. Assessment, Round 1: Fight!


I'm sure some of you are avid readers of ERE (Electronic Recruiters Exchange), but for those of you that aren't (and don't receive my shared items), there's a lively discussion going on over there regarding the practical value of recruitment versus assessment practices.

It started with Wendell Williams' first post on how to identify a bad test (the second part is also worth reading). The comments begin relatively benignly, debating the strength of various predictors of performance (e.g., P-O fit versus behavioral interviews), but turns into, let's say, a lively debate that includes a discussion of the Gallup 12, the limitation of assessments, the Uniform Guidelines, and a lot more. The most heated exchange occurs between Wendell and Lou Adler, where accusations and sarcasm fly.

Speaking of Lou, he continues to advocate his perspective with his next post on whether increasing interview accuracy increases quality of hire (yes, he's suggesting that's an open question). While the comments following are fewer in number, the debate continues regarding the value of assessment and the evidence used to support it (e.g., Schmidt & Hunter's 1998 piece).

Who said HR is boring?

Saturday, February 13, 2010

Latest IJSA: Emotional intelligence, multiple-choice formats, and lots more

The March 2010 issue of the International Journal of Selection and Assessment (IJSA) is out, and the research covers a wide variety of recruitment and assessment topics as well as being truly international:

Unproctored internet-based testing (UIT) response distortion may be less than we fear (sample included cognitive and personality measures)

What factors are most important to organizations when choosing a test? This study suggests applicant reaction, cost, and diffusion of the test type in the field.

Personality (esp. core self evaluation) is related to the type of work preferred, and hence P-O fit

Career site features may differentially attract men and women

Corporate images do matter when it comes to organizational attractiveness

Who uses job-search websites and how to improve them (the sites, not the people)

Support for performance-based (as opposed to self-report) measures of emotional intelligence

Work samples, interviews, and ability tests perceived best by employees (why? because they work, say the participants)

...and last but definitely not least:

A "2 of 5" multiple-choice format seems superior than traditional "1 of 6" (you just have to make sure you can score them that way!)

Sunday, February 07, 2010

Feds new jobs site is Googlish

The U.S. Government has revamped its jobs page, www.usajobs.gov, and in the process shown everyone else how its done.

Take a look at their old site. Not horrible, but cluttered with lots of features that distracted from the main reason people visit the site: to look for a job.


This picture actually doesn't do it (in)justice; there was additional content below the bar.

Now look at their new website:


This new website is what I would call "Googlish": simple, lots of white space, no scrolling required, and a single search box. The design focuses less on being pretty, and more on being functional. If you're interested in learning more about careers, or if you'd like information related to specific groups, like veterans or those with disabilities, its still there. And there's even more functionality up top in the form of drop-down menus.

Job seekers don't need a magazine ad. They need to quickly and easily find information. And this new website fits the bill.

How does yours compare?

Thursday, February 04, 2010

Lessons from NYC Fire case - part 2

Part 2 of 2

Last time I discussed five important lessons we can take away from recent rulings in the Vulcan v. City of New York case. In this post I'll review the remaining lessons and also discuss the relief order.

----

6) The city failed to provide sufficient evidence that the exam(s) tested for a sufficient number of the critical KSAs. They also failed to explain why they chose not to measure several KSAs identified as critical.

Lesson: the courts do not require employers to measure every single critical KSA. But there is an expectation that employers attempt to measure a sufficient number that represent a significant portion of the job requirements. In this case, that included non-cognitive abilities such as resistance to stress, teamwork, and conscientiousness, that were not measured.

7) The city failed to adequately consider how to measure a significant number of essential KSAs. While some of their concerns were valid (e.g., structured interviews for all applicants would be an operational nightmare), there are many different forms of testing that should have been considered, including situational judgment tests (SJTs) and biodata, which can be used to measure non-cognitive components.

Lesson: triers of fact expect employers to be up on the various assessment methods available and be able to explain why they chose not to use certain ones. This includes tests that are relatively easy to develop (e.g., SJTs) as well as ones that require substantial resources and statistical expertise (e.g., biodata).

8) The city failed to conduct a reading level analysis on the exams to ensure that it was not "pointlessly high." The plaintiff introduced evidence suggesting the reading level was above 12th grade; in addition, it appeared to exceed the reading level of materials at the academy.

Lesson: never forget that every assessment method is in some sense measuring additional KSAs beyond those you intend. For written exams, reading comprehension is always a requirement (barring accommodation). It's quite easy to conduct a reading level analysis (MS Word has it built in) to ensure that the level is reasonable and matches other job-related material.

9) The city failed to show that the cutoff scores (pass points) established for the exams were based on adequate rationale, namely "the necessary qualifications for the job of entry-level firefighter." Instead, the cutoff scores were based on operational need (the number of job openings expected). This is particularly important in multiple-hurdle selection processes such as in this case, where a failure on one exam component precludes an applicant from participating in the rest of the (potentially compensatory) assessment process.

Lesson: ultimately applicants have to pass the test(s) to be considered for employment. Cutoff scores should be established using the expertise of both SMEs and test developers and should be based on the minimum competency levels required upon entry to the job. At a minimum (and I would not rely solely upon this), the scores should be analyzed to identify any logical "break-points."

----

After ruling for the plaintiffs on both the adverse impact and disparate treatment claims, the judge issued a relief order on 1/10/10. In it, he imposes several things, including the following:

1) The city must develop a new testing procedure for entry-level firefighter in conjunction with the relevant parties. Following the development of the test, there will be a hearing to determine if this test should be used rather than the current test (developed in 2007 and not at issue in this litigation).

2) The court shall develop a process by which the approximately 7,400 applicants covered by this case can file a claim for monetary relief.

3) The city will identify 293 black candidates on the eligibility list and offer them priority hiring. (No quotas are being imposed, although the judge leaves this possibility open)

4) Retroactive seniority for those hired.

In addition, several other issues are up for debate, including the appointment of a special master or monitor, standards that will be relied upon in constructing the new exam, and the need for additional relief.

---

So what did we learn from all this? If you follow--fairly closely--best practices when developing and administering exams, you will be on solid ground defending them. If you don't, and your exam has a discriminatory effect, you may be called on it--and it's not a pleasant process. I'll leave you with this quote from the January ruling on disparate treatment:

"The history of the City's efforts to remedy its discriminatory firefighter hiring policies can be summarized as follows: 34 years of intransigence and deliberate indifference, bookeneded by identical judicial declarations that the City's hiring policies are illegal."

Sunday, January 31, 2010

Lessons from the NYC Fire case - part 1

Part 1 of 2

New York City, like the cities of New Haven and Chicago, has a long history of employment discrimination litigation related to its firefighter testing.

Since the 1970s and cases like Guardians, the city has been under scrutiny for its woefully low number of black firefighters.

In 2007 the city found itself faced with another lawsuit over its firefighter hiring practices, and in July of 2009, a U.S. District Court judge found that the city had violated Title VII by administering written exams from 1999-2007 that had high levels of adverse impact. The city marshaled an inadequate defense. In January of 2010, the same judge (Nicholas Garaufis) found the city liable for a pattern and practice of disparate treatment for those same exams. An adverse impact finding, particularly for written exams, and especially for public safety tests, is not earth-shattering. But a finding of disparate treatment in this situation is less common.

This case, while only one example and limited in its impact, has some valuable lessons for test users and sheds some light on how judges look at our field. In particular, I describe below nine points the judge specifically made and what lessons we can draw from them:

1) While the city conducted a job analysis with an "extensive" list of tasks and surveyed incumbents, the city offered "no evidence of 'the relationship of abilities to tasks.'" They conducted a linkage, but the judge found that the SMEs were confused about what they were supposed to do and didn't understand several of the abilities they were rating.

Lesson: simply having subject matter experts (SMEs) link essential tasks and knowledge, skills, and abilities (KSAs) is not sufficient. You need to ensure they understand the statements they are linking as well as how exactly they are supposed to be linking them.

2) In conducting the job analysis, the city inappropriately retained tasks and KSAs that could be learned on the job. It is quite clear (e.g., per the Uniform Guidelines) that only tasks and KSAs that are required upon entry to the job should be identified as critical in terms of exam development.

Lesson: make sure that when you are developing exams based on job analysis results that you focus only on those tasks and KSAs that are required upon entry to the job. This should be determined by your SMEs.

3) The city relied to some extent upon the work of a previous test developer, Dr. Frank Landy (who sadly recently passed away). In addition to a tenuous link between Dr. Landy's work and the current exams, the judge makes it clear that "reliance on the stature of a test-maker cannot stand in for a proper showing of validity." At the same time, the judge emphasizes that exams should be constructed by "testing professionals."

Lesson: tests should be developed by people who know what they're doing. This means HR professionals with the requisite background in test validation and construction in conjunction with job experts. Do not rely solely on previous efforts, particularly when (as in this case) the results of those efforts were either incomplete or not fully relevant to your current situation.

4) The city performed no "sample testing" to ensure that the questions were reliable as well as "comprehensible and unambiguous."

Lesson: few steps in the test development process are as easy--or as valuable--as pilot testing. I have yet to see an exam that didn't benefit from a "trial run" with a group of incumbents. Not only will you catch unintended flaws, you will verify that the exam is doing what you claim it is.

5) There was insufficient evidence that the exams actually measured the (nine cognitive) KSAs the city claimed they intended to measure. Plaintiffs were able to suggest the opposite through analyzing convergent and discriminant validity as well as by conducting a factor analysis.

Lesson: there are two linkages of primary importance in test development. The first was describe in #1. The second is the link between critical KSAs and the exam(s). At the very least, you must be able to show evidence that there is a logical link between the two. When you claim to be measuring cognitive abilities, you incur an additional responsibility, which is gathering statistical evidence that supports this claim.

Next time: more lessons and the relief order.

Friday, January 22, 2010

Jan '10 issue of JAP, plus APA gets stingy

The January 2010 issue of the Journal of Applied Psychology is out, and there are some good articles to take a look at. It just may be more difficult to see them. More on that in a minute.

First, here are some of the titles in this issue:

Emotional intelligence: An integrative meta-analysis and cascading model. A must for anyone interested in EI; posits and supports a cascading model whereby emotion perception-->emotion understanding-->emotion regulation-->job performance.

Time is on my side: Time, general mental ability, human capital, and extrinsic career success. GMA shown to have strong links to two extrinsic measures of career success, income and prestige. (ah, but are smart people happier?)

I won’t let you down… or will I? Core self-evaluations, other-orientation, anticipated guilt and gratitude, and job performance. Core self-evaluations' impact on job performance may depend on how much they focus on others.

Understanding performance ratings: Dynamic performance, attributions, and rating purpose. Performance ratings are influences by a variety of things, including overall performance variance and purpose of the ratings.


Okay, so back to my earlier comment: it appears that APA has restricted viewing abstracts of their journals to registered members (hence the lack of links in this post). On the one hand, no big deal, it appears you can simply register to gain access. On the other hand...why should someone have to do this? This is another unfortunate example of research being restricted (first by charging exorbitant fees for articles, now through personal identification) and contributes to the field being insular.

Granted, APA's not the only one that does this (hey buddy, got $400 for the CRL?) but that doesn't excuse it. Our field benefits from sharing of information, not just among professionals but with the general public. Requiring registration does not further that goal. Thankfully some individual researchers (see the sidebar on the main page) allow access to their work--something we should all be grateful for.

Monday, January 18, 2010

How to get r = 1.0


Recruiters have a variety of measures of their success, often including process outcomes (time-to-fill, number of requisitions filled, etc.).

And although assessment professionals have a variety of success measures, some in common with recruiters (e.g., tenure), there is one measure that stands above all others: job performance.

The "gold standard" of this measurement is to correlate test scores with job performance measures (called criterion-related validation evidence). A correlation of, say, .50 between these two, is considered outstanding. Square that and you have the percentage of behavior explained. So in other words, when we can explain 25% of job performance with assessments, we call that success (and with good reason, because it's a heck of a lot better than 0%).

Why not higher than 25%? What would it take to get r =1.0, in other words a perfect correlation between test scores and performance? Here is a somewhat tongue-in-cheek recipe for achieving this impossible dream:

1. An accurate identification of the top competencies/KSAs required for the job. Qualified subject matter experts reach consensus on a handful of far and away the most important qualities that impact job performance.

2. Perfectly constructed and administered, perfectly reliable and accurate measures of the the top KSAs.

3. Variability among applicants in terms of amount of the relevant KSAs possessed.

4. Test scores combined and weighted appropriately given the job analysis results.

5. Variability in scores for those hired.

6. A clear description of the work to be performed and competencies to be demonstrated so the individuals understand expectations.

7. Perfectly reliable, accurate measures of job performance that capture behaviors one would logically relate to the critical KSAs.

8. A supportive work environment (e.g., high quality supervision, adequate resources) so this doesn't interfere with work performance.

9. Variability in job performance among those hired using the assessments.

10. Elimination of outside factors that may contribute to lower job performance (e.g., family emergencies, medical/psychological changes).

As you can see, some of these are achievable (1, 4, 6), others are challenging and depend on circumstances, but are not impossible to achieve (3, 5, 8, 9) and some are practically impossible (2, 7, 10). I said earlier this was tongue-in-cheek because obviously we'll never have a situation where all of these conditions (as well as ones I'm sure I forgot) are true.

Does this mean we should abandon the correlation between test score(s) and job performance? Absolutely not. It should continue to be one of our "gold standards" for measuring our success as assessment professionals. But we--and our customers--should have our eyes wide open before pressing "compute."

Wednesday, January 13, 2010

Meet the new SIOP...same as the old SIOP

The votes are in, and the new name for the Society for Industrial and Organizational Psychology (SIOP) is...the same.

After over a thousand votes from members, the existing acronym beat The Society for Organizational Psychology (TSOP) by a tally of 51% to 49%--a difference of 15 votes. You can read my comments about this option--and my prediction of the outcome--here.

Why is this non-news, news? Because it's problematic that the main professional, scientific body that devotes itself to researching the psychology of organizations and work (POW!) repeatedly has identity issues. This is in large part because of the word "industrial", which makes it sound like we're all studying factory workers. I am not alone in having people look at me sideways when I attempt to explain our field.

To be perfectly honest, I am reluctant to describe my focus as "psychology", except to others in the same field. It sidetracks the conversation (perhaps due to my insufficient skill). It's much easier to connect with people by saying I'm in Human Resources. This isn't to say that the focus on psychology isn't important, or that others in I/O psychology might not mind using this phrase, or that there isn't some brand value in SIOP. But call me crazy, if you're reluctant to name your field (and attorneys don't count--people know, or think they know, what you do), the profession has a problem.

So, our identity struggle continues. An interesting follow-up study might be to ask SIOP members how they describe their field of work to non-I/O folk and break that down by area of focus. It's a big tent.

Personally, I prefer something that includes Work and Organizational. Mix and match letters as you will.

On a positive note, did you know you can access all of SIOP's quarterly news publication, TIP, here? The January 2010 issue has pieces on integrated performance management, a preview of Lewis v. City of Chicago, and a lot more.

Wednesday, December 30, 2009

Outback settlement contains interesting requirements


You may have heard that Outback Steakhouse, a restaurant chain based in Tampa, Florida, has agreed to settle a gender discrimination lawsuit for $19M. What's interesting about this isn't the size of the settlement, but rather the conditions attached.

Background: The EEOC sued Outback in 2006, claiming it systematically discriminated against its female employees by denying them promotion opportunities to the more lucrative profit-sharing management positions. In addition, they claimed that female employees were denied promotional job assignments such as kitchen management, which were required for employees to be considered for top management positions.

The settlement: Outback agreed to a four-year consent decree and $19M in monetary relief. So far, pretty standard. But there were additional settlement requirements, and here's where it gets interesting. In addition to the monetary relief, Outback has agreed to:

1. Create an online application system for employees interested in management positions. This is the first time I've seen this in a settlement (which isn't to say it hasn't happened) and seems to indicate that the EEOC views this as a more "objective" screening mechanism.

2. Create and hire someone for a newly created "human resources executive" position titled Vice President of People. Again, this is a new one for me.

3. Hire an outside consultant for at least two years who will monitor the online application system to ensure women are being provided equal opportunities for promotion and provide reports to the EEOC every 6 months.

The main thing that strikes me about this settlement is the faith that is being placed in an online application system to somehow ensure equal opportunity. Sure, having a standardized application system may cut down on some of the subjectivity of individual hiring supervisors, but it leaves me wondering:

- What will the screening criteria for management positions be?

- How will the outside consultant define "equal opportunities"?

- How will access to the online system be controlled, and who will be making screening/hiring decisions?

- What happens if there continues to be adverse impact, which you would expect if applicants continue to be screened on experience?

- What will be the duties of the Vice President of People, how will they be hired, and how will they interact with the consultant?

This will be interesting to watch.

Sunday, December 20, 2009

Validity: An elusive (unitary?) concept

What makes a test "valid"? What is the best way to develop a selection system? These are two of the most fundamental questions we try to answer as personnel assessment professionals, yet the answers are strangely elusive.

First of all, let's get two myths out of the way: (1) a test is valid or invalid, and (2) there is a single approach to "validating" a test. It is the conclusions drawn from test results that are ultimately judged on their validity, not simply the instruments themselves. You may have the best test of color vision in the world--that doesn't mean it's useful for hiring computer programmers. And many sources of evidence can be used when making the validity determination; this is the so-called "unitary" view of validity described in references like the APA Standards and the SIOP Principles. Unitary in this case refers to validity being a single, multi-faceted concept, not that psychologists agree on the concept of validity--a point we'll come back to shortly.

Although we can debate test validation concepts ad infinitum, the bottom line is we create tests to do one primary thing: help us determine who will perform the best on the job. The validation concept that most closely matches this goal is criterion-related validity: statistical evidence that test scores predict job performance. So we should gather this evidence to show our tests work, right? Here's where things get complicated.

It's likely that many organizations can't, for various reasons, conduct criterion-related validity studies (although baseline evidence of this would be helpful). Most of the time, it's because they lack the statistical know-how or high quality criterion measures (a 3-point appraisal scale won't do it). So in a strange twist of fate, the evidence we are most interested in is the evidence we are least likely to obtain.

So what are organizations to do? Historically the answer is to study the requirements of the job and select/create exams that target the KSAs/competencies required; this matching of test and job is often referred to as "content validity" evidence. But Kevin Murphy, in a recent article in SIOP's journal Industrial and Organizational Psychology, makes an important point: this is good practice, but not a guarantee that our tests will be predictive of job performance. Why not? For a number of reasons, including poor item writing and applicant frame of reference. Murphy makes a passionate argument that we rely way too heavily on unproven content validation approaches when we should focus more on criterion-related validation evidence. Instead of focusing on job-test match, we should focus on selecting proven, high quality exams.

Not surprisingly, the article is accompanied by 12 separate commentaries that argue with various points he makes. It's also interesting to compare this piece with Charley Sproule's recent IPAC monograph where he makes an impassioned defense of content validity.

A complete discussion of the pros and cons of different forms of validation evidence are obviously beyond a simple blog post. My main issues with Murphy's emphasis on criterion-related validation are threefold. First, as stated above, most organizations likely don't have the expertise to gather criterion-related validation evidence for every selection decision (maybe this is his way of creating a need for more I/O psychologists?). Perhaps "insufficient resources" is a poor excuse, particularly for an issue as important as employment, but it is a reality we face.

Second, even if we were to shift our focus to individual test performance, following a content validation approach for development enhances job relatedness (which Murphy acknowledges). Should your selection system face an adverse impact challenge, the ability to show job relatedness will be essential.

Finally, let's not forget that high test-job match gives candidates a realistic job preview--hardly an unimportant consideration. RJPs help candidates decide whether the job would be a good match for their skills and interests. And no employer that I know of enjoys answering this question from candidates: "What does this have to do with the job?"

The approach advocated by Murphy, taken to its extreme, would result in employers focusing exclusively on the performance of particular exams rather than on their content in relation to the job. This seems unwise from a legal as well as face validity perspective.

In the end, as a practitioner, my concern is more with answering the second question I posed at the beginning of this post: What is the best way to develop a selection system? Given everything we know--technically, legally, psychologically--I return to the same advice I've been giving for years: know the job, select or create good tests that relate to KSAs/competencies required on the job, and base your selection decision on the accumulation of test score evidence.

Should researchers work harder to show that job-test content "works" in terms of predicting job performance? Sure. Should employers take criterion-related validation evidence into consideration and work to collect it whenever possible? Absolutely. Will job-test match guarantee a perfect match between test score and job performance? No. But I would argue this approach will work for the vast majority of organizations.

By the way, if you are interested in learning more about the different ways to conceptualize validity--"content validity" in particular--Murphy's focal article as well as the accompanying commentaries are highly recommended. He acknowledges that he is purposely being provocative, and it certainly worked. It's also obvious that our profession has a ways to go before we all agree on what content validity means.

Last point: the first focal article in this issue--about identifying potential--looks to be good as well. Hopefully I'll get around to posting about it, but if not, check it out.

Sunday, December 13, 2009

R/A Predictions for 2010

With 2010 right around the corner, here are some predictions for what the new year will bring in the area of recruitment and assessment:

1) More personality testing. Year after year personality testing continues to be one of the hottest topics. Look for more research, more online personality testing, and new measurement methods.

2) More boring job ads. Even though we know better, don't expect to see any big leaps in readability for 80% of job ads. Same old job descriptions. Maybe we'll see some pictures. On the plus side, more organizations focus on making their career portals attractive.

3) A slow trickle of research on recruiting. The amount of large-scale, sophisticated research on recruiting methods remains a shadow of that found in the assessment literature. Don't expect this to change.

4) More focus on simulations. 2010 sees more focus on simulations, particularly those delivered on-line, as highly predictive assessments as well as realistic job previews. Oh, and they likely have low adverse impact (research, anyone?).

5) Leadership assessment gets even hotter. With the economy improving and more boomers deciding the time is right to retire, finding and placing the right people in leadership positions becomes an even more important strategic objective.

6) Federal oversight agencies get more aggressive. With more funding and backing from the Obama administration, expect to see the EEOC and OFCCP go after employers with renewed vigor. By the way, have you seen the EEOC's new webpage? It's actually quite well done.

7) More fire departments get sued. In the wake of the Ricci decision, fire dept. candidates feel emboldened when they fail a test or fail to get hired/promoted. Look for departments to try to get out ahead of this one by revamping their selection systems.

8) More age discrimination lawsuits. With so many boomers, expect to see more claims of discrimination, particularly over terminations. Keep words like "energetic" and "fresh" out of your job ads.

9) Automation providers slowly focus on simplicity. Whether we're talking applicant tracking or talent management systems, vendors slowly realize that they need to make their applications simpler to increase usability and buy-in. No, simpler than that. Keep going...

10) Employers get more sophisticated about social networking sites. Many realize that rather than jumping on the latest Twitter-wagon, it's best to figure out where these sites fit with their recruitment/assessment strategy. Watch for more positions whose sole role is managing social media.

11) Online candidate-employer matching continues to be a jumbled mess. Without a clear winner in terms of a provider, job seekers are forced to maintain 400 profiles on different sites and may give up altogether and focus more on social networking. Meanwhile, employers continue to try to figure out how to reach passives; LinkedIn continues to look good here but needs to expand its reach a la Facebook.

12) More employers face the disappointing results of online training and experience questionnaires. Will they go back to the drawing board and try to improve them (hint: don't use the same scale throughout), or abandon them for more valid methods, such as biodata, SJT, and simulations? More research on T&Es is badly needed, even if we are just putting lipstick on a pig.

13) Decentralized HR shops centralize. Centralized ones decentralize. Particularly in the public sector, these decisions unfortunately continue to be made based on budgets rather than best practice. Hiring supervisors wonder why HR still can't get it right.

14) Fortunately, HR continues to professionalize. With much of the historical knowledge walking out the door and the job market improving, HR leaders are forced to re-conceptualize how they recruit and train recruitment and assessment professionals. This is a good thing, as it means more focus on analytical and consultative skills.

Keep up the good work everybody. And Happy Holidays!

Sunday, December 06, 2009

Setting cutoff scores on personality tests


What's the best way to set a cutoff score for a personality test, knowing that some candidates inflate their score? It all depends on your goal. Are you trying to maximize validity or minimize the impact of inflation?

According to a research study by Berry & Sackett published in the Winter '09 issue of Personnel Psychology, if your goal is to maximize validity, your best bet is to wait until applicants have taken the exam, then set your cut-score (e.g., the top two-thirds); this was particularly true when selection ratios are small (i.e., organization is very selective).

If your goal is to minimize the number of deserving applicants who are displaced by "fakers", you're better off establishing the cut point ahead of time, by using a non-applicant derived sample (e.g., job incumbents, research group). The results were generated using a Monte Carlo simulation.

Interestingly, the authors also replicated the work of other researchers who have shown that the impact of faking on the criterion-related validity of personality measures is relatively low. There are a few other very good points made in this article:

- Expert judgment methods of establishing pass points (e.g., Angoff method) may be difficult to use for personality tests since experts may find it difficult to judge individual items. Methods used to select a certain number of applicants or methods based on a criterion-related validity study (both used as variables in this study) are more appropriate for personality tests.

- There is no consensus of how prevalent faking on personality exams is; estimates range from 5-71%. It likely depends on the situation and how motivated test takers are to engage in impression management.

- Some recommend setting a very low cutoff score for personality tests, which would exclude only those likely not suitable for the position (and not faking), while others prefer a more stringent cutoff to maximize utility.

- A reasonable range of d-values for score inflation on personality inventories is .5-1.0 (used in this study).

- There exists very little research on the skewness of faking score increases. A positively-skewed distribution (meaning most people faked a small amount) was used in this study. (I would think this would also vary on the situation)

So bottom line: where--and how--you set your cutoff score on personality inventories depends on whether you want to maximize the predictive validity or minimize the number of deserving applicants that get left out of the process.

Other good reads in this issue:

- Police officer applicants reactions to promotional assessment methods

- The impact of diversity climate on retail store sales

- The construct validity of multisource performance ratings

- Labor market influences on CEO compensation

Tuesday, November 24, 2009

Want better prediction? Gather more data.


That's the bottom line from a study in the November 2009 issue of the Journal of Applied Psychology.

Oh & Berry looked at how adding personality ratings from peers and supervisors added incremental validity to self-ratings using a five-factor model measure. What were the results? Increases of 50-74% in operational validity across personality facets. They also looked at differential prediction of task and contextual performance (unfortunately those results weren't reported in the abstract). Bottom line? If you're using a personality assessment for promotions, strongly consider gathering data from co-workers.

Speaking of self-presentation, in the same issue Barrick et al. report the results of a meta-analysis of how self-presentation tactics (e.g., appearance, non-verbal behavior) impact interview ratings and later job performance. Results? "What you see in the interview may not be what you get on the job and...the unstructured interview is particularly impacted by these self-presentation tactics." An important reminder of how who the candidate seems to be impacts your assessment, and another reason to collect multiple sources of data.

There are a number of other great articles in this issue, such as:

How Major League Baseball CEO personalities impact important outcomes (like, um, winning).

How SJT and biodata measures add to the prediction of college student performance.

How personality scale validities change over time among a group of medical students.

Differences among letters of recommendation in academia between genders.

Saturday, November 21, 2009

HR: C-level or AAA?

Just when we thought we were C-level, Dogbert reminds us how old-school CEOs view HR.

p.s. I think I like CPO better than HCO.

Friday, November 13, 2009

Explain away

Of all the low-hanging fruit in recruitment and selection, perhaps none are easier to implement than explaining your process. It's shocking how few selection processes are fully (and coherently) explained, not just in terms of getting from point A to point B, but why the points exist in the first place.

Turns out explaining the process to candidates matters. A lot. And while we already knew that applicant perceptions were important, a recent meta-analysis by Truxillo, et al. published in the International Journal of Selection and Testing clarifies the impact that explanations have. Specifically, explanations were related to:

- Fairness perceptions (important in their own right)

- Perceptions of the hiring organization

- Test-taking motivation

- Performance on cognitive ability tests

Furthermore, fairness effects were greater when paired with personality tests rather than cognitive ability exams.

What does all this mean? It's pretty simple, really--communicate, communicate, communicate. Explain in clear terms to all applicants what the full selection process is, and why. Imagine something like this:

"Thank you for your interest in applying for the Blog Reader position. The selection process will consist of the following:

Step 1. You submit your application and work sample by December 5.

Step 2. Your application is reviewed to ensure you meet our minimum qualifications found on the job posting.

Step 3. If so, your work sample is scored by internal subject matter experts. It will be judged on relevance to the position, complexity, and contribution to the profession. The top scorers move on to Step 4.

Step 4. You are contacted by Human Resources to set up an interview. The interview will last approximately 2 hours and will take place at our San Jose campus. They will take place in mid-January.

Step 5. You interview with the hiring supervisor and 2-3 potential co-workers. You will be asked a series of questions designed to measure your knowledge of Blogs. You will also be asked to complete a writing sample which will be judged for style as well as content. Top performers in these steps will be asked for a final interview with the head of the Blogger Division.

Should you fail to proceed to any Step, you will be contacted and told why."

That wasn't so hard, was it?

For more about applicant perspectives in selection, check out the entire December '09 issue.

And while I'm on the subject of research, check out the latest issue of the International Journal of Testing for good stuff on test compromise, DIF, and P-O fit. Oh, and don't forget about the entire issue devoted to test adaptation.

Last but not least, Practical Assessment, Research, and Evaluation (PARE) has put out several things lately worth reading, check 'em out.

Tuesday, November 10, 2009

What employers can learn from Twilight


Not since Harry Potter have I seen such obsessed fans. The buzz started several months ago in anticipation. And November 20th is almost here.

Don't know what happens on November 20th? Ask your daughter. Or granddaughter. Or, heck, pretty much any woman between the age of 16 and 30. That's the day that New Moon, the second installment in the wildly popular Twilight series, hits theaters. Why should employers care about this, other than anticipating that certain staff members will be out of the office that day? Read on.

Like Harry Potter, the Twilight books (written by Stephenie Meyer) are enormously popular, and the movies are too. Also like Harry Potter, fans range the demographic spectrum, although the most rapid fans seem to be women (not surprising given the protagonist and the love triangle she's in the middle of).

Most employers would kill to have the kind of brand devotion that Twilight fans have. If Twilight was an employer, they would be competing with Google for top talent. So what can we learn from this phenomenon that can help us with branding our organization?

1. People like good stories. Twilight was a phenomenon way before Robert Pattinson and Taylor Lautner (the male lead actors). Whether you're in banking, IT, public utilities, or a flower shop, you have stories to tell about how your organization has impacted others--or how your employees have impacted each other. Are you telling these stories, or letting stories be told about you?

2. People still read. Related to #1, the Twilight phenomenon, like Harry Potter, began with the books. There's a lot of hype about video these days, but given something interesting, people have no problem spending time reading it. What does your recruitment material look like--is it entertaining? Educational? Would you read it even if you weren't interested in a job there?

3. People like fantasy. There's an awful lot of reality out there right now--the recession, H1N1, wars--and people like to take a mental break. Don't be afraid to break out of the mold and try telling a story that takes people away from their day-to-day lives.

4. People like contests and "sides". One of the biggest, possibly the biggest, dramas within the world of Twilight is the competition between the two main lead male characters. Fans identify themselves as being on "Team Edward" or "Team Jacob". This isn't something you see employers do very often, and it requires a bit of built-in loyalty, but it's something that can engage fans even more.

5. People like being fans. There's something primal about being part of a group of people who share the same interest. If you give people something to be a fan of, they'll enjoy connecting with others who share their passion. This is what the Facebook fan pages are all about. Google has 300k+ Facebook fans. Twilight has over 4 million.

6. You can brand almost anything. Branding is about more than your website, or your recruitment fliers. It can become part of everything your organization does, if it's strong enough. And it's more than just a logo, it's about the organization's philosophy and accomplishments. Look around and you can probably see Twilight branded on almost everything. I'm surprised they don't have Twilight adhesive bandages. Oh wait, they do.

As November 20th approaches, be prepared for a media onslaught about Twilight. Whether you're a fan or not, use it as an opportunity to think about how your organization could garner that kind of excitement. After all, that's what leads high potentials to want to apply.