.comment-link {margin-left:.6em;}

Sometimes I Wish That It Would Rain Here

Friday, February 22, 2008

evaluation, interpretability, and the utility of dreaming

a few nights ago, my girlfriend and I watched The Science of Sleep. for those not familiar, the basic premise revolves around a main character who has difficulty discerning between his waking life and his dreams. afterward, we got into (what I felt was) a somewhat confused conversation about whether or not it was a good movie, why it was a good movie, and how you know it's a good movie (and I fear that conversation has led indirectly to what is an almost equally confused blog post). one of the pseudo-conclusions to which we came is that it's a good movie because of its interpretability. that is, there are several parts of the movie, foremost the ending but also bits and pieces throughout, where it was, we think, left intentionally unclear what exactly happened. the point is not to figure out the "true" or "real" story at those points. rather, the genius of the movie seemed its ability to engage the audience in interpreting those somewhat ambiguous parts. my mind kept slipping towards questions of evaluation; how do we know it's a good movie? when the key aspect of the movie has nothing to do with the movie objectively and everything to do with the interactions between movie and viewer, how can one really say anything about the movie itself?

evaluation is a huge buzz word in HCI. "ok, cool, you built your system, but does it work? does it achieve the intended goal?" I've heard it described that part of a dissertation is scoping out a problem, picking a portion of that problem, describing the win condition wherein you know that the problem has been solved, and, crucially, demonstrating that the win condition has been achieved. even when we recognize that evaluation is an interactive process, that it's really about the meeting of system and user, evaluation so often boils down to a question of success. does the system achieve the goals it set out to accomplish? there are certainly lots of conversations going on right now about richer, fuller means of evaluation, focusing less on system evaluation and more on experience evaluation, and emphasizing that evaluation is a process of determining value, which is necessarily contextually (historically, culturally, etc.) contingent. personally, I find a lot of this work both particularly compelling and very liberating, especially with respect to the epistemological questions it raises; what do we as a field consider valid knowledge, and how do we validate methods of knowledge production? on the other hand (maybe it's just that I'm having a hard time shedding my positivist roots), I have a desire to know, does it work?

this desire becomes inherently problematic when the ostensible goal of a system is to support, facilitate, encourage, and even engender critical thinking and reflection, especially when that reflection hinges on the interpretability of the system. here, I refer to interpretability not as a question of "do participants interpret this system properly?" rather, the question I want to ask is, "to what extent does the system present a resource for interpretation?" it's difficult enough to ascertain whether or not interactors are engaging in these abstract process--critical thinking, reflection, interpretation--to begin with. now, try to determine to what extent the interactor's thoughts, feelings, behavior are a result of interacting with the system. the very notion seems misguided; we're not dealing with a system cause-and-effect relationship here, but rather a whole complex system in which I doubt any single aspect can be causally linked to any other. besides, this isn't about controlling for confounding factors. it's about getting people to think, to critically engage, and to question, reconsider, and possibly even reformulate their conceptual frame.

I think one of the difficulties in my case is that the system I'm developing has a goal that seems objectively evaluable. does it do what I say it does? am I able to automatically identify conceptual metaphors (a la Lakoff and Johnson) in bodies of written text? well, I think the question of whether or not it works, or how well it works, depends largely upon the interactor's (i.e., "user's") interpretation of the system's results. moreover, I think it hinges on the interpretability of those results. the question, I suspect, should not be, "does the system accurately and correctly identify conceptual metaphors?" rather, the question should be, "does the system produce results that serve as a resource for the interactor's interpretation, and through that interpretive process does the interactor engage in critical thinking and reflection?" not that this is a particularly easy question to answer, but it seems a somewhat more useful one in terms of evaluating, i.e., determining the value, of the system. it's not about measuring success, it's about understanding the interactors' experience with the system.

this ended up getting too long for a single post, so I'll end with the above thoughts about evaluating interpretability and reflection. more stuff about dreaming to follow...

Labels: , , ,

Wednesday, February 06, 2008

iSchool dissertations and rigor

I was following this discussion at UW's iSchool, and the previous week's, on what constitutes an information science dissertation, and I found the (posted) results of this panel rather interesting. they list a series of pragmatic suggestions from students, the first and most important being to "satisfy yourself first," followed by a list of expectations from professors. One striking expectation from the professors is that the dissertation be "rigorous." I suspect it is no coincidence that this expectation is followed by being able to justify a qualitative dissertation to a quantitative researcher and vice versa. While we can all agree that research should be rigorous, what actually constitutes rigor seems to vary, at times greatly, and not just along quant/qual lines. The faculty in the department at my school seem similar to UW: folks from different disciplines coming together due to common interests. This leads to different definitions of what counts as rigorous, sort of a panoply of perspectives from which to choose the best fit for your particular problem.

I'm wondering, though, as iSchools begin to graduate students, what will the field of information science consider rigorous? will it maintain this panoply approach? will certain approaches get canonized and others become discredited? will we generate new approaches distinct to iSchool-type problems? what are the potential ramifications if these methods get picked up and transfered to other fields? when and how might such new methods be considered rigorous, both in info sci and in other disciplines?

Labels: , , , ,

Sunday, May 06, 2007

CHI - experience evaluation SIG

last week was CHI, and there are about a bajillion things that came up that I really want to blog about. the first one that's actually made it out of my head and onto my keyboard was about the SIG I went to on evaluating experience-based HCI, organized by Joseph 'Jofish' Kaye, Kirsten Boehner, Jarmo Laaksolahti, and Anna Ståhl. due to delayed Caltrains, I didn't get there until a half hour into the SIG, so I missed pretty much all of the introductions. however, it was a nenjoyable, exciting thing of which to be a part, and I feel like there was some truly useful discussion. the participants at the SIG broke up into 3 small groups, each of which chose to address one particular question. my group (including, among others, Michael Muller, Janet Vertesi, Ryan Aipperspach, and Sara Ljungblad) took on, "What are good criteria for an evaluation of experience-focused HCI?" essentially, this is a question of meta-evaluation, that is, how do we evaluate our methods for evaluating experience-focused HCI. the first bit below is a number of criteria that our group thought would be good for evaluating evaluation methods. below that is just a transcription of my notes from the SIG. most of these are from little comments jotted in the margins of the sheet of paper they gave us, so they come in no particular order. the regular type face are my actual notes, and the italics are comments I added after the fact.

What are good criteria for an effective evaluation method?

- Does it highlight the role of the researcher? of the observed?

- Does it elicit rich stories? think descriptions? (I think there are interesting problems with elicitation here that tie back into the first point)

- Does it help to construct a faithful analysis or account or report?

- Does it inspire users/designers/researchers/companies?

- Does it open up interpretations? (closing out isn’t necessarily bad)

- Does it make sure not to generate graphs?

historicizing objectivity – Daston and Galison (Representations, 40, 81-128) describe how objectivity means different things at different times

standpoint epistemology and The Voice from Nowhere

experiences in the moment vs after the moment. how to get people back into the moment after the fact? can use reflective visualizations. rather people to discuss around an artifact

highlight the reflexive nature of technology (is/can technology itself be reflexive? reflective?)

challenges and opportunities for subjectivity

expose subjectivity

the process of an experience vs the product of an experience (I suspect this might be a distinction between having an experience and the memory of the experience. hm, is the act of remembering an experience itself, possibly quite different and distinct from the original experience being remembered? in experience-focused HCI/design, perhaps we could/should support not only having experiences but also the experience of remembering those experiences)

evaluating experiences as storytelling, plumbing a collection of episodic memories (the notion of stories and narratives came up a lot in our small group)

elicit multiple, different stories. different experiences for different users. retain multiple persectives. (emic perspective, multiple realities. there’s almost certainly a connection here to something a whole lot bigger about multiple thought styles, epistemological pluralism, different cognitive framings, etc.)

literary theory often evaluates texts over and over. why do we not return repeatedly to the same UX to understand it more fully or in different ways? (perhaps this connects to the above point, in that HCI doesn’t particularly value having lots of different perspectives, so studying the same thing again is seen as a waste of time)

often talk about the “representative” user. how is a particular user representative? statistically? politically (e.g., elected union representative)? who chooses the representative and how?

different criteria to evaluate different evaluation methods for different experiences

evaluation as developing a sensibility rather than determining progress – “the tyranny of progress” (a distinction that came up in the HCI and New Media workshop in which I participated was that of evaluation vs analysis, that evaluation might be more along the lines of judging something as good or bad, while analysis is more about understand something. those perhaps might not be the right terms, but I suspect it might be an important distinction to make)

process of presenting results of experience evaluation: exposure -> awareness -> empathy -> advocacy -> change

a member of another group said “when we’re in flow, we’re having an experience,” that flow can be one indicator of an experience (this raises the interesting notion that we might sometimes be having an experience and might at other times not. I’m not sure how much I’d agree with this notion that we might at some times not be having an experience, but I’d certainly agree that we have different types of experiences at different times, and that when trying to evaluate experience you might only be interested in certain types of experiences.)



I'd love to hear others' thoughts/comments/questions about this stuff. I feel like it's an important direction for HCI to pursue, and I think that having discussions about how to pursue it most effectively is an important aspect of making experience-based evaluation more accepted and central to the HCI community.

Labels: , , , ,