Computerworld
Quick Menu
Search



Ads by TechWords

See your link here


Subscribe to our e-mail newsletters
For more info on a specific newsletter, click the title. Details will be displayed in a new window.
Data Management
Computerworld Daily News (First Look and Wrap-Up)
Computerworld Blogs Newsletter
The Weekly Top 10
More E-Mail Newsletters 
Computerworld 2007Subscribe to Computerworld
40 years of the most authoritative source of news and information for IT leaders.

QuickStudy: Web Harvesting

 

Sign up to receive Data Mining Resource Alerts

June 21, 2004 (Computerworld) -- It's hard to argue with the proposition that the World Wide Web is the largest repository of information that has ever existed. In just over a decade, the Web has moved from a university curiosity to a fundamental research, marketing and communications vehicle that impinges upon the everyday life of most people in the developed world. But there's a catch, of course. As the amount of information on the Web grows, that information becomes ever harder to keep track of and use.


More
Computerworld
QuickStudies


This vast amount of freely available information is spread over billions of Web pages, each with its own independent structure and format. So how do you find the information you're looking for in a useful format—and do it quickly and easily without breaking the bank?

Search Isn't Enough

Search engines are a big help, but they can do only part of the work, and they are hard-pressed to keep up with daily changes. For all the power of Google and its kin, all that search engines can do is locate information and point to it. They go only two or three levels deep into a Web site to find information and then return URLs. They also find and return meta descriptions and meta keywords embedded in Web pages, but these may well be inaccurate.

Consider that even when you use a search engine to locate data, you still have to do the following tasks to capture the information you need:

  • Scan the content until you find the information.
  • Mark the information (usually by highlighting with a mouse).
  • Switch to another application (such as a spreadsheet, database or word processor).
  • Paste the information into that application.

A better solution, especially for companies that are aiming to exploit a broad swath of data about markets or competitors, lies with Web harvesting tools.

Web harvesting software automatically extracts information from the Web and picks up where search engines leave off, doing the work the search engine can't. Extraction tools automate the reading, copying and pasting necessary to collect information for analysis, and they have proved useful for pulling together information on competitors, prices and financial data of all types.

Continued...
1 | 2 | NEXT  



Print this Story Send Us Feedback E-mail this Story Digg! Digg this Story Slashdot this Story
Sidebar: Web harvesting and libraries
Sidebar: Resources for more information on Web harvesting
Web Harvesting
Sidebar: Also Known As ...
"Yahoo's owners have spoken and Jerry Yang is out as Yahoo CEO. Does this mean that Microsoft is in?..." Read more...
"In five years, IT will still be a viable career path, but it will no longer exist as the stand-alone..." Read more...
Read more Business Intelligence posts or See all Blogs
Microsoft dissed Intel's 915 chipset before making 'Vista Capable' changes
Microsoft dumps OneCare, slates free security software for '09
Ballmer: Yahoo acquisition won't happen, despite Yang's departure
More top stories...
Obama administration to inherit tough cybersecurity challenges
BlackBerry Storm sales should be strong, Verizon says
Google deal produces 91% of Mozilla's revenue
If you're like our 7,000 survey respondents, your paycheck this year has been flattened and your bonus obliterated. We offer 12 ways to plump up your paycheck.
Microsoft's next OS might more accurately be called Windows 6.5: It's essentially a better version of Vista.
Twitter can be a valuable business tool -- if you know what you're doing. Here's how to juice it for all it's worth.
By helping Intel with loosened 'Vista Capable' requirements, Microsoft 'severely damaged' its credibility, said an HP exec in a newly unsealed Feb. 2006 e-mail.
Get the latest news, reviews and more about Microsoft's newest desktop operating system
Find wage data for 50 IT job titles.
All Zones
Business Continuity Zone
The File Data Management Zone
Security Management Zone
The SAS Zone
Business Intelligence and Analytics Zone
The Enterprise Search Zone
Software as a Service Zone
The Security Zone

Ads by TechWords

See your link here
Speeding the time to intelligence
Get this Computerworld report free for a limited time, compliments of SAS.
Time To Intelligence -- a concept defining how long it takes to get accurate and timely information into the hands of workers who need it most. Do it slower than your competitors and your company is toast. Do it faster, you scorch them. Business Intelligence is the key to optimizing Time To Intelligence, and success there is a combination of people, policies, and technology.
Download this executive briefing download
Transforming Disaster Recovery - VMware Infrastructure for rapid, reliable and cost-effective Disaster Recovery
Download this white paper today!
(Source: VMware) VMware Infrastructure transforms disaster recovery by providing you fast, reliable and cost-effective disaster recovery. Why suffer from the slow, expensive and unreliable problems associated with traditional disaster recovery solution? VMware makes disaster recovery affordable through consolidation savings and re-use of existing servers for your disaster recovery site. Experience the speed of virtualization!
Download this white paper go
Turning information into a Competitive Advantage
Turning information into a Competitive Advantage
View this webcast now!
Go to the webcast 
White Papers
Read up on the latest ideas and technologies from companies that sell hardware, software and services.
Deploying Virtualized NetWare on Linux Whitepaper
Collaboration Tools and Organizational Success
Driving Business Success Through Workgroup Choice and Flexibility
View more whitepapers 

SAS Information Management Kit

SAS is the leader in business intelligence and analytical software and services. Only SAS offers leading data integration, storage, analytics and business intelligence applications within a comprehensive enterprise intelligence platform. SAS gives 97 of the top 100 companies in the 2007 Fortune 500 THE POWER TO KNOW®.

Webcast: The Information Management Roadmap
Imagine high-quality data, cleansed, analyzed and delivered throughout your organization. Join Computerworld, IT visionary Thornton May and a panel of experts to learn how SAS® can help you make it happen.

View this webcast 
Research Report: Information Management Initiatives at Midsize and Large Organizations
See the top-line results of this Computerworld sponsored survey to see how IT and business leaders are handling information management implementation.

Download this report 
White Paper: Information Management: Better Information for Winning Decisions.
This white paper explains how the SAS Information Evolution Model aids companies in assessing how they use this information to make strategic decisions and drive business.

Download this white paper