- Hive — open source crowdsourcing framework from NYT Labs.
- Prezi Data Team Culture — good docs on logging, metrics, etc. The vision is a great place to start.
- Scroll Behaviour Across the Web (Chartbeat) — nobody reads above the fold, they immediately scroll.
- threat_research (github) — shared raw data and stats from honeypots.
More than algorithms, companies gain access to models that incorporate ideas generated by teams of data scientists
Data scientists were among the earliest and most enthusiastic users of crowdsourcing services. Lukas Biewald noted in a recent talk that one of the reasons he started CrowdFlower was that as a data scientist he got frustrated with having to create training sets for many of the problems he faced. More recently, companies have been experimenting with active learning (humans1 take care of uncertain cases, models handle the routine ones). Along those lines, Adam Marcus described in detail how Locu uses Crowdsourcing services to perform structured extraction (converting semi/unstructured data into structured data).
Another area where crowdsourcing is popping up is feature engineering and feature discovery. Experienced data scientists will attest that generating features is as (if not more) important than choice of algorithm. Startup CrowdAnalytix uses public/open data sets to help companies enhance their analytic models. The company has access to several thousand data scientists spread across 50 countries and counts a major social network among its customers. Its current focus is on providing “enterprise risk quantification services to Fortune 1000 companies”.
CrowdAnalytix breaks up projects in two phases: feature engineering and modeling. During the feature engineering phase, data scientists are presented with a problem (independent variable(s)) and are asked to propose features (predictors) and brief explanations for why they might prove useful. A panel of judges evaluate2 features based on the accompanying evidence and explanations. Typically 100+ teams enter this phase of the project, and 30+ teams propose reasonable features.
Laura Busche looks at trends through a Lean Startup lens
Your code can be clean as a whistle and your software deployment on-time, but if you’re a startup, branding is as vital to your success as any non-crashing app. The following trends are discussed through a frame of Lean Startup practices, relevant to any startup.
I sat down with the one-in-a-million Laura Busche, author of the upcoming book Lean Branding, at the recent Lean Startup Conference in San Francisco. We talked about what she sees as important branding trends for 2014. Those trends follow below, in both text and video formats.
You’ll find one additional surprise trend covered below, the video portion of which you can see in the Top Lean Branding Trends for 2014 compilation video below.
Are there any patterns within the nine trends discussed? Considered as a whole, a feeling of intimacy and intrigue certainly stand out. The intimacy element comes in the form of big brands trying to let customers feel closer to them by acting like smallish local brands, as well as the intimacy of working directly with customers to create the message of a brand (crowdsourcing) rather than lecturing potential customers about the merits of a brand. The intrigue element comes in the form of the use of images rather than words to communicate brands, playing on immediate emotional response, and also from brands surprising customers with their unusual personalities and intentional quirkiness, as well as by turning up in customer conversations to solve problems.
AI Book, Science Superstars, Engineering Ethics, and Crowdsourced Science
- Society of Mind — Marvin Minsky’s book now Creative-Commons licensed.
- Collaboration, Stars, and the Changing Organization of Science: Evidence from Evolutionary Biology — The concentration of research output is declining at the department level but increasing at the individual level. […] We speculate that this may be due to changing patterns of collaboration, perhaps caused by the rising burden of knowledge and the falling cost of communication, both of which increase the returns to collaboration. Indeed, we report evidence that the propensity to collaborate is rising over time. (via Sciblogs)
- As Engineers, We Must Consider the Ethical Implications of our Work (The Guardian) — applies to coders and designers as well.
- Eyewire — a game to crowdsource the mapping of 3D structure of neurons.
Zombie Drones, Algebra Through Code, Data Toolkit, and Crowdsourcing Antibiotic Discovery
- Skyjack — drone that takes over other drones. Welcome to the Malware of Things.
- Bootstrap World — a curricular module for students ages 12-16, which teaches algebraic and geometric concepts through computer programming. (via Esther Wojicki)
- Harvest — open source BSD-licensed toolkit for building web applications for integrating, discovering, and reporting data. Designed for biomedical data first. (via Mozilla Science Lab)
- Project ILIAD — crowdsourced antibiotic discovery.