Wednesday, August 20, 2008

Remind me again...

...why we're supposed to trust governments rather than corporations with our private data?

Monday, August 18, 2008

Hunh

Some scientists have done a study of the forwarding patterns of Internet chain letters and come up with two interesting findings:

1) Ninety percent of the time, when someone forwards a chain letter to a group of people, only one person out of that group will forward it on.

2) At the median, when a person receives an Internet chain letter it will have been forwarded three hundred times before it reaches them.

I'm not sure what any of this means yet, but I find it intriguing. I'm particularly curious how this compares to pure information, as opposed to chain letters. I can't imagine that a piece of information would go through three hundred episodes of telephone before it reaches a person . . . would it?

Full paper here.

Wednesday, July 30, 2008

The Things One Can Discover...

...with enough data and the right algorithm:

Falling Coca-Cola sales in a specific region of Africa are an excellent indicator of civil unrest, famine, or some other problem in that region.

Sunday, July 13, 2008

Now That's What I'm Talking About!

A system that analyzes the content of books and suggests new books for you to read based on the amounts of dialog and action, the density, the pacing, and a couple of other factors. The books currently in the beta are heavily skewed towards science fiction, which is a genre I haven't really read in since I was 16, so I can't say how well it works based on what's in there now. But kudos for the general concept!

The Slashdot story has links to more background on the project and the person behind it.

Friday, July 4, 2008

Be a Data Geek, Win a Prize!

The British government is going to give 20,000 pounds to the person who comes up with the best idea for mashing up and re-using the reams of data collected by said government.

I checked the rules, and it's not limited to British citizens. The deadline is in September. Full details here.

Sunday, June 29, 2008

Skewz

It took me awhile to decide whether this new Skewz site is brilliant or one of the signs of the apocalypse, but I think I've come down on the side of "brilliant."

Basically, the point of Skewz is to use the wisdom of crowds to make explicit the bias that exists implicitly in the media, while also functioning sort of like Digg, aggregating stories that people find interesting. People submit stories and then get to vote on how much and in which direction they think the stories are skewed.

But Skewz isn't designed to let people only see the stories that they agree with ideologically—the pages are divided up into one column for "liberally skewed" stories and one for "conservatively skewed" stories, which lets readers see both sides of the news right next to each other. Unfortunately Skewz doesn't seem to actively pair conservative and liberal stories on the same topic, but still, it's useful to be able to see what both sides are talking about on any given day—especially since quite a bit of media bias isn't in how a story is covered but in what people think is newsworthy in the first place.

And, one of the best parts: it aggregates people's skew ratings for each story and uses them to create a giant chart showing the ideological skew of quite a few of the major newspapers and blogs on 20 different issues.

So of course this comes out after I quit editing the books where I had to pair opposing viewpoints and was always running around trying to find somebody arguing the other side of some mildly obscure issue....

Hat tip: Marginal Revolution (again!)

Friday, June 13, 2008

Well That's New and Interesting

While doing a search in Google today I got the following line at the bottom of the first page of results:

"In response to a complaint we received under the US Digital Millennium Copyright Act, we have removed 1 result(s) from this page. If you wish, you may read the DMCA complaint that caused the removal(s) at ChillingEffects.org."


I wonder if Google has been doing this for awhile and it just so happened that today was the first time I happened to do a search that hit one of these, or if this is a new thing?

Saturday, May 31, 2008

There Has Got to Be an Easier Way

Tyler Cowen at Marginal Revolution has a post up on how he finds new books to read. (If you're not reading Marginal Revolution yet—and you should be!—Prof. Cowen reads scarily massive numbers of books on wildly divergent topics.) Here's just a partial list of the things he does to find new books (bracketed expansions of acronyms are mine): “visit Borders every Tuesday to look for new books, go to a local public library every other day and scan the new books section, subscribe to TLS [Times Literary Supplement], London Review of Books, New York Review of Books, noting that you should spend more time with the ads than the book reviews, read the blogs Bookslut and Literary Saloon, read the new magazine BookMark (recommended), read the NYT [New York Times], FT [Financial Times], and Guardian and their books sections....” (FYI, he's a professor of economics, not English, so it's not like keeping up on new fiction is part of his job description.)

This reminded me of a post I've been meaning to write for awhile about the person I know who had the most trouble finding new books to read—my grandmother. Well, no, let me rephrase that slightly. She outsourced the job of finding books for her to read to my mother and me, so really it was us who had the book-finding problem. My grandmother went through a book about every day and a half, so every two weeks, when my mother and I went to the library, we had to find about 10 books for her. And she was a picky reader. She liked love stories best, but only if there was no sex or bad language in them; she would read lighthearted mysteries, like the Mrs. Pollifax and Cat Who series, but nothing violent or dark or scary.... It was really all that my mother and I (and the wonderful librarians at the Middletown Public Library) could do to keep her in books. Our saving grace was that, once three or so years had passed, she would forget that she had read a book, so we could give it to her again.

How did we keep track of which books she had already read? At that time the library kept a card in the back of each book with the library card numbers of everyone who had checked out the book, along with their respective due dates. So we just had to look for our library card number in the list and see how long it had been since we'd last checked that book out. "Horrors!," I can hear the librarians out there thinking. "Freedom to read! It is unethical for the library to keep records like that of who has checked out a specific book! What if the government subpoenas those records?" I understand the logic behind that position (although I also don't think that my grandmother really would have cared if the government found out about her love of the novels of Janette Oake), but at the same time, I can't even imagine how we would have kept my grandmother in books without those records. We lived out in the boondocks, so it wasn't like we could make a mid-week run to the library and get more books for her if we accidentally brought home a pile of books that she had read recently. The library was too small for us to give her exclusively new books, and ILL was no solution—in the dark days before Amazon, it was hard to get much information about the content of a book without holding it in your hands. (Library of Congress subject headings don't really tell you things like, "How graphic are the murders in this murder mystery?" or "Is there out-of-wedlock sex in this romance novel?")

So when I say that libraries really ought to consider doing something Amazon-like to help people find books that they might like, I'm thinking of all of the time that my mother and I spent over the years flipping through books to decide if my grandmother might like them or not and scrutinizing the little card in the back to see how recently she had read them. If you estimate that we spent half an hour every two weeks doing that for probably thirty years (well, I only participated in this process for about ten years, but I think my mother did it for close to thirty)...that's a lot of time that could have been saved if the library catalog had had some system for saying, "People who like the kind of books that you like also like these new books." And that's one of the ethical values of librarianship too, right? Saving the time of the reader?

(Oh, and for another argument for why libraries ought to be taking a page from Amazon's book, check out the comments to Prof. Cowen's post. Count the number of people who advise using various Amazon features to get good book recommendations, and the number who know enough about how those methods work to recommend ways to improve the Amazon recommendations. Now count the number of people who say anything at all about libraries/librarians as a way of finding new books. By my count, the ratio is 7 to 0.)

Friday, May 9, 2008

Another Benefit of E-books/E-articles/etc.

They're much easier to move.

I'm just about finished packing. Just out of curiosity, I counted the boxes of information that I'm moving—not counting documents that I need to keep for legal or record-keeping reasons or anything like that, just books and papers that I'm keeping purely for the value of the information in them.

The tally:


  • 4 milk-crate-sized boxes of notebooks/papers/printed articles/photocopies/etc. from my undergraduate and graduate courses

  • 1 milk-crate-sized box of coursepacks from my undergraduate courses

  • 17 smallish boxes of books

  • A half-box of cookbooks

  • Two years' worth of American Libraries and Information Technology in Libraries



I hereby resolve to think about moving 22-boxes-plus worth of paper next time I'm tempted to print an article to write notes on it or to buy a paper book that I could get as a (not DRMed-to-death) e-book. Henceforth (or at least until I'm settled into someplace that I have no intention of moving out of ever again) all of my information is going to be electronic.

Thursday, May 1, 2008

Don't Trust Everything You Read in Books, Part 274

A new biography of King Louis XIV's mistress, Madame de Maintenon, got the whole way through the editorial process at Bloomsbury (one of the more prestigious British publishing houses) without anybody noticing that one of its sources, a “diary” supposedly written by the king, was actually historical fiction.

Stories like this (as well as 6.5 years of working in the publishing industry) are a big part of why I worry that traditional information literacy instruction does a disservice to students by encouraging them to rely on external authority cues (Was it published by a reputable publisher and/or in a peer-reviewed journal?) rather than on internal accuracy cues (Do their numbers add up? Can you track down and verify their sources?) when evaluating information. Not everything that's made it through the editorial process is true, and not everything on the Web is false, and it seems like students would be much better served by learning how to evaluate the truth of the message rather than the “trustworthiness” of the medium.

Sunday, April 27, 2008

More on Twine

I'm still playing with Twine, and it's growing on me. It's like del.icio.us on steroids.

So, I've started a twine called The Examined Web, where I'm going to collect stories about the social, economic and political impacts of Web 2.0 / the Semantic Web / [insert other Web-related buzzwords here]. That means that I probably won't be doing Link Roundups here anymore, unless I've got a group of stories that I want to comment on and not just point people to. So if you're interested in that kind of stuff, join the twine! You have to join Twine itself first, but I've got 10 invitations for it, so leave a comment here or e-mail me and I will make sure you get one if you need it.

And by the way, if you're interested in the technical side of the Semantic Web and you're planning on joining Twine, you should check out Apps :: On Semantic Web & Related Applications. They've got some great stuff.

(Update: links to the twines added, although I'm not sure what happens if you click on those links and you haven't joined Twine yet....)

Friday, April 25, 2008

The Cult of the Amateur Blogger Makes a Comeback!

I know I've argued that blogs deserve more respect than they get from some people, but even I think that this is going a bit too far.

The Cult of the Amateur Blogger No More?

Some people have been arguing that blogging is going to kill traditional journalism because free bloggers will undercut paid journalists.

Today, Megan McArdle points out that most of the good amateur bloggers have now been hired by one media corporation or another.

Anecdotes != data, but it's an interesting anecdote none the less.

Wednesday, April 23, 2008

Finally!

Six months ago I posted about how Twine was going to make my life easier, once I got my beta invitation.

Well, it finally came today, and I've spent the past hour and a half poking around in it.

It's definitely still in beta (real beta, not Google's "beta-in-perpetuity" beta), and the algorithms they're using to pull metadata out of free text still need some work, but I can definitely see the potential in it. But I'm not sure that I see as much potential as the stories from 6 months ago were promising. For what I'd primarily be using it for (organizing my personal research/bookmarks), Zotero has it beat by a mile at this point—and that will jump to about 10 miles once Zotero gets around to launching the server sync and recommender services that they're promising.

Ah well. I will probably play around with it a little further, as it moves out of beta into something that's actually supposed to be fully functional, to see how it winds up.

Monday, April 14, 2008

Access to Information in the Third World

There's a nice article in this weekend's New York Times Magazine about how enhanced access to information improves living conditions in the third world. "Ah hah," I'm sure all of you librarians out there are now thinking, "Further proof of how wonderful libraries are!" Actually, no: the story is about how cellphones allow the global poor easy access to information that was either completely unavailable or prohibitively expensive before. (It's an interesting mental exercise for a librarian, actually, to read this article and try to figure out how a library could function to meet the sorts of information needs that are featured in it.)

By the way, the article mentions in passing the story of the fishermen of Kerala, India, who provided some of the first evidence of how important access to information is for the global poor. That story by itself is fascinating. If you're interested, here are a paper from the Quarterly Journal of Economics and a Washington Post article about the research that has been done on these fishermen.

(Yes, I am aware that it's been almost a month since I updated this blog. Yes, I am aware that this makes me a bad blogger. Blogging will become more regular once I finish packing up all of the junk I've acquired in the past 6.5 years and moving it 500 miles.)

Monday, March 17, 2008

Another Advertising-Related Link Roundup

There's been a big dust-up in Britain over the past couple of weeks about targeted advertising. As somebody who thinks that advertising-supported content is a good thing (and yes, I still do at some point intend to do a long and thoughtful post about why I think that), I don't necessarily find all of the arguments against targeted advertising to be persuasive. But I still think they're interesting and worth listening to.

So, here are a few of the better entries in the debate:

A Liberal Democratic politician says, in relation to targeted advertising on MySpace, “I think it's absolutely wrong if you haven't been notified and given the opportunity to opt out.” My take: Notification is definitely a good thing, but opt-out is a very different question. MySpace can only afford to give you a free account because advertisers are willing to give them money to show you advertisements. Why should you be able to say to MySpace, “I want you to give me my free account, but I refuse to help you make the money to pay for that account”? It seems to me that the opt-out option is, “If you don't like MySpace's advertising practices, don't use MySpace.”

I don't know much about the details of this Phorm system, but it's interesting that there's one story praising its privacy-protecting features, another a few days later saying that there isn't enough information to know if it's acceptable privacy-wise or not, and another one saying it's illegal. (More on Phorm.) And Tim Berners-Lee has come out against not only Phorm, but all systems that track online activity in order to provide targeted advertisements.

Friday, March 14, 2008

Yahoo! Search Is Going Semantic

Or so the Yahoo! Search Blog said yesterday, anyway. Details are still a little thin, but Dublin Core is first on the list of metadata schemas that will be supported....

Friday, March 7, 2008

More on Privacy, Data and the Government

Apropos of the comments I made a couple of weeks ago about privacy and the government, there's a nice article in Wired right now that addresses some of the same issues.

(Hat tip: Slashdot, where the comments are actually pretty interesting too.)

Tuesday, March 4, 2008

How Did I Not Know about This?

There is an entire blog dedicated to quirky numeric data, with a strong emphasis on visualizations thereof.

I am in heaven.

Thanks to Marginal Revolution for pointing this blog out and making my day.

Saturday, March 1, 2008

Link Roundup

The Atlantic has a very interesting article on Internet censorship in China. (And there's also a Web-only interview with the author.)

The Encyclopedia of Life, an open encyclopedia that aims eventually to have comprehensive entries on every single species of living thing, has launched.