Registration is now open! Early-bird rates until Feb. 20.
Hotel reservations are open too! (off-site link) Conference rates until Feb. 20.
Financial aid is available! Application deadline is Feb. 11.
Dates to remember:
* Tutorials: Thursday, March 13, 2008
* Conference: Friday, March 14, through Sunday, March 16, 2008
* Sprints: Monday, March 17 through Thursday, March 20, 2008
The Python Software Foundation is proud to present PyCon 2008, the 6th annual Python community conference, this year in Chicago.
Come to PyCon 2008 in Chicago to:
* meet interesting people in the Python community
* learn about cool things others have done recently
* show off the cool things you've done recently
* learn about projects, tools & techniques
* advance open source projects
There's a lot happening at PyCon:
*
Tutorial Day: Thursday March 13
A chance for in-depth learning from the experts.
*
Conference Days: Friday March 14 to Sunday March 16.
o Keynote talks from luminaries.
o Scheduled talks on a myriad of subjects.
o Lightning talks: 5-minute impromptu talks, scheduled at the conference (come prepared to talk!).
o Open Space: plenty of rooms available for follow-up discussions after scheduled talks, group discussions of projects, unscheduled talks, demos, brainstorming...
o And the hallway track is always popular!
*
Development Sprints: Monday March 17 to Thursday March 20.
Up to four days of intensive learning about and development of open-source projects. We provide the space and the infrastructure (network, power, tables & chairs), you provide the enthusiasm and the brainpower. All experience levels welcome. The sprints are free!
*
Expo Hall: Friday March 14 & Saturday March 15.
Many sponsors and vendors will have booths in an expanded exhibitor area this year. Get a T-shirt, find a job, talk to the companies who participate in the Python community!
But most importantly,
*
The Community: PyCon is a community conference!
PyCon is organized and run by volunteers from the Python community. Please join us!
The conference venue is the Crowne Plaza Chicago O'Hare hotel. Reservations should be made here (off-site) in order to receive the conference rate.
Latest News (PyCon Blog)Subscribe to the RSS Feed
# Reminder: Financial Aid Application Deadline is Feb. 11th
2008-02-08 22:24:30
Just a quick reminder -- applications for financial aid for PyCon 2008 are due on Monday, February 11th! Please see the PyCon website for application ...
# Intro to Sprints sessions
2008-02-06 17:02:27
There will be two pre-sprint sessions that will be run on Sunday afternoon, after the end of the conference talks.Intro to Sprinting (60 minutes, at 3...
# Publicizing PyCon
2008-01-29 14:27:53
We need you to let people know about PyCon!PyCon has always relied on the community to get the word out. This year, we've put together a "Publicizing...
# PyCon 2008 Financial Aid Available
2008-01-22 20:07:29
The Python Software Foundation has allocated some funds to help people attend PyCon 2008. If you would like to come to PyCon but can't afford it, the ...
# PyCon 2008 Registration Open!
2008-01-21 01:17:12
I am pleased to announce that PyCon 2008 registration is now open! Early-bird registration is open until February 20, so there's one month to registe...
Monday, February 11, 2008
Sunday, February 10, 2008
Is Microsoft Office Adware?
Posted by Andrew Z at Saturday, February 9, 2008
Is Microsoft Office adware? Wikipedia defines adware as "any software package which automatically plays, displays, or downloads advertising material to a computer after the software is installed on it or while the application is being used."
In Microsoft Office Professional 2003's help, a search for "APA" (a popular documentation style) brings up two links labeled Microsoft Office Marketplace.
Microsoft Office Word 2003 help showing two marketplace ads
A click on the link opens a web page where the item is available for a fee. Also, there's a banner advertisement from Dell.
Screenshot of Microsoft Office Marketplace
A search in Office's help for "print" leads to the brief article "Print more than one copy" with two up-selling links for Office 2007.
Microsoft Office Word 2003 help showing two up-selling ads for Office 2007
By the way, clicking either promotion launches Internet Explorer even when Firefox is the default browser. The reason is the Office help runs an embedded Internet Explorer showing a page from office.microsoft.com. Despite the technical explanation, it can be confusing and inconvenient for the user to use a non-default browser.
On Microsoft.com, Sandi Hardmeier, MVP, concludes her adware definition, "Ads are not bad by themselves but they become a problem when they are unauthorized. Unfortunately, many adware programs do not give users enough notice or control." In Office, where is the "notice or control"? A workaround is to search the Offline Help instead of the default Microsoft Office Online.
Could Office be spyware? Microsoft defines spyware as "software that performs certain behaviors such as advertising." Microsoft continues:
That does not mean all software that provides ads or tracks your online activities is bad. For example, you might sign up for a free music service, but you "pay" for the service by agreeing to receive targeted ads. If you understand the terms and agree to them, you may have decided that it is a fair tradeoff. You might also agree to let the company track your online activities to determine which ads to show you.
Basically, spyware includes adware, but not all ads are bad. Are ads bad about after paying for an Office license?
Another part of spyware is tracking. The Office 2003 EULA (.PDF and .XPS wrapped in .EXE) states:
CONSENT TO USE OF DATA. You agree that Microsoft and its affiliates may collect and use technical information gathered as part of the product support services provided to you, if any, related to the Software. Microsoft may use this information solely to improve our products or to provide customized services or technologies to you and will not disclose this information in a form that personally identifies you.
While in Office the collected information is not personal, the EULA does not mention disabling collection.
The Office Privacy Statement admits the use of tracking cookies. While cookies are normal and generally harmless, it is unusual to require cookies or to use them in a desktop application. If cookies are disabled in Internet Explorer, the Office help fails.
Microsoft Office Word 2003 help: an error when cookies are disabled in Internet Explorer
In conclusion, Office 2003 does display ads, and certain parts require cookies. While these are a normal and healthy part of the web, it is, at least, unusual for a commercial desktop application.
Is Microsoft Office adware? Wikipedia defines adware as "any software package which automatically plays, displays, or downloads advertising material to a computer after the software is installed on it or while the application is being used."
In Microsoft Office Professional 2003's help, a search for "APA" (a popular documentation style) brings up two links labeled Microsoft Office Marketplace.
Microsoft Office Word 2003 help showing two marketplace ads
A click on the link opens a web page where the item is available for a fee. Also, there's a banner advertisement from Dell.
Screenshot of Microsoft Office Marketplace
A search in Office's help for "print" leads to the brief article "Print more than one copy" with two up-selling links for Office 2007.
Microsoft Office Word 2003 help showing two up-selling ads for Office 2007
By the way, clicking either promotion launches Internet Explorer even when Firefox is the default browser. The reason is the Office help runs an embedded Internet Explorer showing a page from office.microsoft.com. Despite the technical explanation, it can be confusing and inconvenient for the user to use a non-default browser.
On Microsoft.com, Sandi Hardmeier, MVP, concludes her adware definition, "Ads are not bad by themselves but they become a problem when they are unauthorized. Unfortunately, many adware programs do not give users enough notice or control." In Office, where is the "notice or control"? A workaround is to search the Offline Help instead of the default Microsoft Office Online.
Could Office be spyware? Microsoft defines spyware as "software that performs certain behaviors such as advertising." Microsoft continues:
That does not mean all software that provides ads or tracks your online activities is bad. For example, you might sign up for a free music service, but you "pay" for the service by agreeing to receive targeted ads. If you understand the terms and agree to them, you may have decided that it is a fair tradeoff. You might also agree to let the company track your online activities to determine which ads to show you.
Basically, spyware includes adware, but not all ads are bad. Are ads bad about after paying for an Office license?
Another part of spyware is tracking. The Office 2003 EULA (.PDF and .XPS wrapped in .EXE) states:
CONSENT TO USE OF DATA. You agree that Microsoft and its affiliates may collect and use technical information gathered as part of the product support services provided to you, if any, related to the Software. Microsoft may use this information solely to improve our products or to provide customized services or technologies to you and will not disclose this information in a form that personally identifies you.
While in Office the collected information is not personal, the EULA does not mention disabling collection.
The Office Privacy Statement admits the use of tracking cookies. While cookies are normal and generally harmless, it is unusual to require cookies or to use them in a desktop application. If cookies are disabled in Internet Explorer, the Office help fails.
Microsoft Office Word 2003 help: an error when cookies are disabled in Internet Explorer
In conclusion, Office 2003 does display ads, and certain parts require cookies. While these are a normal and healthy part of the web, it is, at least, unusual for a commercial desktop application.
Waving the flag: NetBSD developers speak about version 4.0
By Federico Biancuzzi | Published: January 30, 2008 - 11:53PM CT (Whole Article)
Introduction
The NetBSD community announced last month the official release of NetBSD 4.0, the latest version of the Unix-like open-source operating system. Version 4.0 includes significant new features like Bluetooth support, version 3 of the Xen virtual machine monitor, new device drivers, and improvements to the Veriexec file integrity subsystem. NetBSD, which is known for its high portability, is capable of running on 54 different system architectures and is suitable for use on a wide range of hardware, including desktops, servers, mobile devices, and even kitchen toasters.
Meet the developers
To commemorate the NetBSD 4.0 launch, enthusiast Federico Biancuzzi communicated with 21 developers to produce this expansive interview with loads of insightful information about the NetBSD 4.0 development process.
NetBSD Foundation secretary Christos Zoulas will discuss why sendmail has been removed and what goals the project has for the fundraising process. He wrote a lot of code, such as svr4 emulation, isapnp code, ptyfs, siginfo, ELF loader, mach-o loader, statvfs, pam integration, etc.
Liam J. Foy will explain the delay in the release engineering process.
Elad Efrat's area of interest in NetBSD is mostly enabling security technologies. From allowing flexible fine-grained security policies on the system, to actually writing them and making them easy to deploy, he is interested in minimizing the time it takes to construct a secure installation. His major contributions to NetBSD, if trying to chronologically order them, are Veriexec, fileassoc(9), kauth(9) (and the secmodel(9) derivatives), security model abstraction (bsd44, securelevel), PaX features (MPROTECT, Segvguard, ASLR), pw_policy(3). These topics were summarized in a paper published on SecurityFocus in late 2006 and later presented in EuroBSDCon 2006. Subsets were also presented in smaller academic and military venues in Israel.
Matt Fleming will describe how veriexecgen works.
Nicolas Joly will present the status of Linux binary compatibility.
Matthias Scheler is currently mostly working on "pkgsrc" and occasionally fixing bugs in the base system. He was the responsible release engineer who managed the NetBSD 3.0 release. He will talk about the X Window System and the new digital transfer mode in cdplay.
Joerg Sonnenberger contributed a lot of infrastructure work for pkgsrc and he is the responsible for modular Xorg in pkgsrc. His contributions in the NetBSD base system are ACPI and x86 improvements.
Jason Thorpe will describe proplib(3).
Manuel Bouyer is a NetBSD user since 0.8. His first contribution was porting the OpenBSD shared library support for NetBSD/pmax in 1995, then he started working on ATAPI devices support. In late 1996/early 1997, he also wrote the ext2fs support, starting from the ffs code (no GPL code in it). He got invited to becode a developer at this time to integrate his work on ATAPI and ext2fs. After that, he kept working in the ATA and SCSI areas (he added the primary DMA support for PCI ATA controllers for example). On occasion, he also worked on other parts of the kernel (Ethernet controller drivers, of dual-endian support to FFS, for example).
When Xen 2.0 was out, he started working on domain0 support so that he could build virtual servers not relying on Linux. He added support for Xen 3.0 when it got out, and became NetBSD/xen portmaster at about this time. He also became a releng member in early 2007 to give some manpower to the team. He will give an overview of NetBSD/Xen and the new 'no-emulation' eltorito boot method.
Phil Nelson installed and ran the first web server serving www.netbsd.org. (the service has since been taken over by project servers.) He still does some minor admin work for the web server.
He runs the WWU build cluster and improved the scripts for the build clusters so they could build releases faster. He wrote a program called xarcmp that compares the contents of archives without having to extract them. It is part of the standard tool now used by both build clusters.
He wrote the initial versions of menuc, msgc and sysinst. sysinst is the current install system and it uses menuc and msgc. He will talk about LFS (Log-structured File System).
Julio M. Merino Vidal has been a member of The NetBSD Foundation and a developer since November 2002. He started contributing to the pkgsrc project with the main goal of getting GNOME 2 to work under NetBSD. He then made other contributions to the core NetBSD operating system, the most relevant of which are the tmpfs file system and the automated testing framework (ATF). He will discuss the features of tmpfs.
Antti Kantee worked on pkgsrc, device drivers and hardware support for various platforms, file systems, and assorted kernel and userland bits. He maintains the file(1) utility in the NetBSD source tree and NetHack in pkgsrc. He will describe the relationship between puffs, FUSE, and ReFUSE.
Alistair Crooks is the president of The NetBSD Foundation, Core Team member, and founder of pkgsrc.
He contributed to pkgsrc, user management software - user(8), ReFUSE (BSD-licensed re-implementation of FUSE), numerous filesystems based on ReFUSE, and iSCSI target and initiator. He will make clear which parts of the iSCSI protocol are included in this release and how much "hackathons" are helpful.
Reinoud Zandijk will add details about the implementation of the Universal Disk Format (UDF).
Iain Hibbert's major contribution has been the Bluetooth protocol stack and associated drivers and utilities. He talks about Bluetooth support in NetBSD.
Arnaud Lacombe has been a developer for about one year now. At the beginning, he joined the project to fix bugs found by the Coverity scan of NetBSD. He is interested in porting NetBSD to new embedded platforms. OpenMoko was a really interesting way to do it, but he never managed to get hardware, so this project is stalled for the moment.
Martin Fouts, Noud de Brouwer, and Arnaud discuss the current support for the hardware included in the iPhone and OpenMoko (Neo1973).
Alan Ritter will comment on the possibility of using the NDIS wrapper to port binary drivers among platforms.
Yamamoto Takashi will describe what agr(4) is.
Jan Schaumann used to work a lot on the web site. He also ported pkgsrc to IRIX and did bulk-builds there. He is a member of the communication-exec team and managed NetBSD's participation in the first two Google Summer of Code. He will discuss their experience with the Google Summer of Code project.
Release engineering, Sendmail, and kauth
What happened to the Release Engineering process for 4.0?
Liam J. Foy: Basically, the release engineering was started as planned. However, after the release engineering started a lot of changes were made which would be too time consuming to pull up (pull up from current to the NetBSD 4branch). Thus, we started the process again with all the changes merged.
You may remember a few NetBSD hackathons which took place. Well, these are what caused the large number of changes.
Why did you remove sendmail?
Christos Zoulas: Sendmail has been, is, and will be a security accident waiting to happen (unless it is rewritten from the ground up with security consciousness). Performing character pointer gymnastics in 50-100 line loops does not create any warm and fuzzy feelings for me. To top this off, most sendmail security issues are marked as confidential, and we are prevented from fixing or mentioning the problem until the ban is lifted. The last time this happened, we said "enough" and removed it altogether before the ban for that particular security issue was lifted
Would you like to present kauth, the kernel authorization framework, included in this release?
Elad Efrat: Kernel authorization is really something cool that I'm happy many people will get a chance to test in this release.
Basically, the story goes like this: the traditional Unix security model, the one where root is almighty and everyone else is not, is slowly beginning to show signs of age as demands for finer-grained security policies arise. The problem becomes more apparent when you look at the kernel code and see that, probably as time went by, people used various ways to check whether a certain operation is allowed to be performed or not—most are variants of checking whether the effective user-id is 0 or using the suser() function—effectively embedding the security model to the kernel, in a way that the question "Can this user open a raw socket?" really became "Is this user root?"
Kauth(9) was originally designed by Apple, and it provides a dispatching model that allows us to really ask if a certain operation can be performed by specifying it along with the relevant context. This abstracts the security model from the various kernel subsystems. The system is divided to "scopes" that collect actions of the same nature (right now, NetBSD has scopes like "process", "network","machdep", etc.), and each scope has a set of "actions" that define the operation (the context provided depends on the operation). When the kernel wants to ask if a certain operation is allowed, it calls the authorization wrapper for the relevant scope, providing it the action and context, and receives back a binary response of "yes" or "no".
This upper layer, like I said, effectively abstracts the way these decisions are made, allowing one to "plug" (almost) any security model one can think of.
These security models are implemented by attaching "listeners" (kauth(9) terminology for callback functions) that receive the action and context and make the decision. In the future it is hoped that these listeners will be able to run on an entirely different machine, thus allowing a centralized security policy for an entire network to be controlled from a single host.
Unfortunately, in NetBSD 4.0, although the security model abstraction is almost complete, there is still not full control over all operations. That is—in some places, the question is still "Is this user root?", so implementing different security models will not be complete. However, it will provide an indication of how well kauth(9) works and how users react to this new ability to develop custom security models (for example, classic uses like restricting raw networking, or binding to certain low ports, to a few users). I'm hoping to be surprised by our user base's creativity.
I highly recommend spending a few minutes reading kauth(9) and secmodel(9) in the current man pages. There's a little taste of what's in the plans for NetBSD 5.0—like credential inheritance control—with the ultimate goal being to provide system administrators with security policy control that requires either zero or very little effort.
PaX, fileassoc, and Veriexec
What security features have you added to mprotect(2) from PaX?
Elad Efrat: First, the PaX project, for those who aren't familiar with it, is responsible for all the modern exploit mitigation technologies. It's where stuff like W^X and ASLR (Address Space Layout Randomization) were born, along with many, many other cool features. One of them, which I initially ported to NetBSD, is PaX MPROTECT.
PaX MPROTECT can be thought of as "strict W^X". If in a normal system, a program starts with no memory pages that are both writable and executable, but those pages can still be created using mprotect(2), systems or programs that run with PaX MPROTECT are "immune" to attacks where protection on pages is modified, often by trashing arguments to mprotect(2).
paxctl(8) is a tool that enables PaX MPROTECT on a per-program basis, if you don't want to enable it globally. I'm afraid that not much testing was done on various architectures using this feature, so users should first experiment with it...
How does the new fileassoc KPI (Kernel Programming Interface) work?
Elad Efrat: I'll begin by describing what fileassoc(9) is.
Some (newer) filesystems allow on to attach metadata to files by using extended attributes (in addition to the usual ones—permissions, timestamps, etc.) and such. A common use for them may be, for example, to store ACLs (Access Control List) for the file that allow finer granularity on access control.
The problem is that these extended attributes are filesystem-dependent and either don't exist on most filesystems or are interfaced differently. This is where fileassoc(9) comes into play: it allows you to attach metadata to files in a filesystem-independent way while storing the information in fast-access kernel tables. The advantage is that it really is filesystem-independent, so there's the potential of adding features and such to filesystems that lack them (for example, again, ACLs). The disadvantages are that the metadata is not really "attached" to the file in any way, and must be loaded in some form—from a database file or such—to the kernel. So that the dependency becomes on the OS itself, but that's cool, because it's aimed for NetBSDsystems.
The KPI works by storing data in hash tables. Metadata is identified by a key (a C string) and is a stream of bytes. A kernel subsystem can then use the fileassoc(9) KPI to either attach, query, or remove metadata from files, by specifying the key.
The fileassoc(9) KPI is implemented in the VFS (Virtual File System) layer, identifying each file using its "file handle" (these are supposed to be unique). So pretty much any filesystem is supported. An example (and the only, at the moment) consumer of fileassoc(9) is Veriexec, where a database file holds information that is parsed by a userland tool, then fed to fileassoc(9) using a special device.
Plans to write a generic interface for communicating with fileassoc(9) have come up in the past, but since we're not yet convinced of the necessity, these are just plans at the moment. Developers interested in feeding data to fileassoc(9) should also implement their own special device stub—at least for now.
Is there any news about Veriexec?
Elad Efrat: As with NetBSD 4.0, I feel Veriexec has gotten to a point where it should probably be used by most of our users.
Veriexec is NetBSD's integrity subsystem, which, in short, can guarantee the integrity of the programs, configuration files, etc., on the system.
In addition to the tons of improvements in performance and stability, and the many features added, NetBSD 4.0 introduces 'veriexecgen', which is a tool written by Matt Fleming (mjf@). Veriexecgen tremendously lowers the bar for fingerprint database generation to a point where it's possible to run a single command with no arguments after installation to have Veriexec set-up appropriately. I strongly recommend reading the veriexecgen(8) main page and giving it a try.
What is veriexecgen and how does it work?
Matt Fleming: veriexecgen is a program that runs against directories and generates a set of fingerprints for use with veriexec. This fingerprint database is usually stored in /etc/signatures.
XFree86, pkgsrc, proplib, and Xen
Do you still include Linux binary compatibility?
Nicolas Joly: I'm currently working on improving Linux binary compatibility on AMD64 for both 32- and 64-bit applications.
The amd64 compat linux is the first one that include NPTL (NativePOSIX Thread Library) emulation found on Linux 2.6 kernel. Likewise, compat linux32 is pretty new, and even if mostly identical to i386linux emulation needs to be modified.
For NetBSD 4.0, kernel support for compat linux is not enabled bydefault as this is not as stable as other ports. For this release, do not expect to run complicated linux applications yet, but basic ones should work. In the mean time, -current has made some progresses...
Which X Window System is included in NetBSD 4.0?
Matthias Scheler: XFree86 4.5.0.
The XFree86 Project has unfortunately lost a lot of momentum. NetBSD is currently in the process of switching to the X.org X11 distribution. We initially tried integrating the monolithic X.org distribution into our xsrc source tree. But it never reached a state where it worked on all platforms and supported cross builds via build.sh.
After the X.org project changed to a modular distribution, it was obvious that pkgsrc is the best way to integrate X11 in the future. There is ongoing work mostly by Joerg Sonnenberger to make modular X11 in pkgsrc cross-buildable. When that work has been finished and properly integrated with the system installation, the XFree86 distribution in xsrc will probably be retired.
What's new in pkgsrc?
Joerg Sonnenberger: I'm only aware of NetBSD 4 being a requirement to use the cross-compiling support. The updates to pkg_install will be in NetBSD4.1, so that is ruled out.
The cross-compiling support allows a small subset of pkgsrc to be built for any architecture running NetBSD using the output of build.sh. Currently this subset includes modular Xorg and a few applications like Xpdf as proof of concept. This is intended to replace the aging XFree86in xsrc, but needs some more polishing work on the pkgsrc side.
What is proplib(3) and what can we use it for?
Jason Thorpe: proplib is a library for manipulating property lists. Property lists are collections of properties, typically stored in a dictionary. A dictionary is an associative array, essentially key-valuepairs. The keys are strings, and the values are strings, numbers, opaque data, booleans, arrays, dictionaries, ...
This is handy in a variety of applications. For example, communicating structured data between userspace and the kernel, describing properties of devices, etc.
What can we do with NetBSD/Xen?
Manuel Bouyer: Well, we can do a lot of things, from merging different physical boxes to a single one (if you need different OSes for different tasks, for example), to easily create test systems (it's very convenient for kernel developement: a guest boots much faster than a regular PC). With the virtualization features of recent X86 CPUs (which are supported by Xen3 with NetBSD as dom0), you can also boot plain i386 OSes. This can be used to test install media, for example, or to load systems that can't be paravirtualized (e.g., Windows).
Filesystems
What is the status of LFS (Log-structured File System)?
Phil Nelson: LFS is in use on about 1/2 of the 23 nodes of the WWU build cluster. Granted, it is running an older 4.0-Beta kernel, but they have been running just fine for quite a while. Currently all the machines in the WWU build cluster have been up 78+ days and continuously building. They are all i386 machines.
Would you like to describe the features of tmpfs?
Julio M. Merino Vidal: tmpfs was born as a replacement for MFS—the memory file system included in all BSDs as far as I know—and as such it shares many functionality with it. It basically is an efficient, memory-based file system which means that it uses part of the system's virtual memory to store files in a way that is more space- and speed-efficient than MFS. It supports all of the features expected from a Unix-style filesystem, including hard links, symbolic links, devices, permissions, file flags and NFS exportability, but it currently does not support sparse files.
As opposed to MFS, tmpfs file systems can grow and shrink automatically depending on the available memory if they are configured as such; their memory consumption can also be upper-limited. And even though the end user cannot directly notice it, tmpfs's code is much simpler than MFS's one, which means that it is easier to audit, test and optimize.
As you just mentioned, tmpfs is a bit different than MFS, but why did MFS need a replacement in the first place?
Julio M. Merino Vidal: The main problem of MFS is that it performs poorly, uses too much memory, and the used memory cannot be reclaimed unless the file system is unmounted. To understand why this happens, we need to see how MFS works, and to see that, we first need to outline how FFS is designed.
FFS (or UFS), the traditional BSD file system, is conceptually split in two different layers. the upper layer handles all the logic from the file system, including the management of directories, the common file operations, the routines to map data to blocks, etc. The lower layer lays out the data on disk, allocating blocks and inodes as needed, etc. This seems like a very nice design, but in the end, it imposes restrictions on the way the lower layer operates (or at least that's the impression I gathered; I'm not too familiar on FFS's code).
MFS is just a replacement of FFS's lower layer to operate on virtual memory. The approach it takes is dead simple: it allocates a contiguous block of memory and treats that as if it were a disk, organizing the contents of such memory region as disk blocks. As you can expect, such approach has a lot of overhead because the file system uses an incorrect abstraction—disk blocks and cylinder groups—to manage memory.
tmpfs, on the other hand, is a much simpler filesystem which takes advantage of the fact that it is always memory-backed. It uses traditional structures, linked lists, arrays, memory pools and other sorts of data abstractions to represent the contents of the file system. These abstractions are easier to deal with and hence require less resources to manage.
It is also interesting to mention that tmpfs grew a custom regression testing suite that was also well received by developers. But the way it was created was suboptimal, which made me start the ATF (Automated Testing Framework) project this year. We'll see a better testing suite in NetBSD 5.0, or at least I hope so! Oh, and by the way, FreeBSD also imported tmpfs in its source tree.
What is the relationship between puffs, FUSE and ReFUSE?
Antti Kantee: puffs stands for Pass-to-Userspace Framework File System. It is an interface and a framework implementation for userspace filesystems on NetBSD. The userspace filesystem interface itself is heavily influenced by the kernel virtual filesystem interface. puffs can be thought to consist of two parts: the mechanism for transporting file system requests from the kernel to userspace and a userspace library, libpuffs, for interfacing with the kernel and writing the userspace filesystem implementation.
FUSE, or Filesystem in Userspace, is another API for implementing userspace filesystems. It is native to Linux, but has been widely ported to other operating systems such as FreeBSD and OpenSolaris. FUSE is considered the standard interface for implementing userspace filesystems with numerous filesystems readily available.
ReFUSE is the implementation of the FUSE API for NetBSD. It is implemented on top of puffs and completely done in userspace. Other non-Linux operating systems implement FUSE compatibility in the kernel, but we believe the ReFUSE approach is the right way to address the issue: export the kernel filesystem interface in the most natural way possible (puffs) and implement compatibility in userspace (ReFUSE).
A very important difference to note is API stability. Since the FUSE API is well-established, it is stable, and filesystems written for it will most likely continue to function for a long time. The puffs API is a "low-level" API which follows the NetBSD kernel virtual filesystem API fairly closely. As the virtual filesystem API on NetBSD evolves, so will the puffs API (and vice versa, if things go as planned). A benefit from this evolutionary incompatibility is the ability to use all the features provided by the kernel API at the best possible performance.
What is their status in NetBSD 4.0?
Antti Kantee: puffs is still under heavy development. The version of puffs found in 4.0 is a snapshot of what was in the tree at the time 4.0 was branched. This means that ReFUSE and therefore FUSE support is unfortunately not present. The most useful application for users is likely to be ssshfs, simple sshfs, which can be found in source form from the tree under src/share/examples/puffs/ssshfs. It implements sshfs functionality. This implementation was superceded by mount_psshfs(8) for NetBSD 5.0.
Technically there is no reason why support for FUSE could not be added to the NetBSD 4 branch. However, it requires backporting ReFUSE and some puffs features from the development branch and therefore a person with time and motivation to do the work, and, above all, test and support it.
Introduction
The NetBSD community announced last month the official release of NetBSD 4.0, the latest version of the Unix-like open-source operating system. Version 4.0 includes significant new features like Bluetooth support, version 3 of the Xen virtual machine monitor, new device drivers, and improvements to the Veriexec file integrity subsystem. NetBSD, which is known for its high portability, is capable of running on 54 different system architectures and is suitable for use on a wide range of hardware, including desktops, servers, mobile devices, and even kitchen toasters.
Meet the developers
To commemorate the NetBSD 4.0 launch, enthusiast Federico Biancuzzi communicated with 21 developers to produce this expansive interview with loads of insightful information about the NetBSD 4.0 development process.
NetBSD Foundation secretary Christos Zoulas will discuss why sendmail has been removed and what goals the project has for the fundraising process. He wrote a lot of code, such as svr4 emulation, isapnp code, ptyfs, siginfo, ELF loader, mach-o loader, statvfs, pam integration, etc.
Liam J. Foy will explain the delay in the release engineering process.
Elad Efrat's area of interest in NetBSD is mostly enabling security technologies. From allowing flexible fine-grained security policies on the system, to actually writing them and making them easy to deploy, he is interested in minimizing the time it takes to construct a secure installation. His major contributions to NetBSD, if trying to chronologically order them, are Veriexec, fileassoc(9), kauth(9) (and the secmodel(9) derivatives), security model abstraction (bsd44, securelevel), PaX features (MPROTECT, Segvguard, ASLR), pw_policy(3). These topics were summarized in a paper published on SecurityFocus in late 2006 and later presented in EuroBSDCon 2006. Subsets were also presented in smaller academic and military venues in Israel.
Matt Fleming will describe how veriexecgen works.
Nicolas Joly will present the status of Linux binary compatibility.
Matthias Scheler is currently mostly working on "pkgsrc" and occasionally fixing bugs in the base system. He was the responsible release engineer who managed the NetBSD 3.0 release. He will talk about the X Window System and the new digital transfer mode in cdplay.
Joerg Sonnenberger contributed a lot of infrastructure work for pkgsrc and he is the responsible for modular Xorg in pkgsrc. His contributions in the NetBSD base system are ACPI and x86 improvements.
Jason Thorpe will describe proplib(3).
Manuel Bouyer is a NetBSD user since 0.8. His first contribution was porting the OpenBSD shared library support for NetBSD/pmax in 1995, then he started working on ATAPI devices support. In late 1996/early 1997, he also wrote the ext2fs support, starting from the ffs code (no GPL code in it). He got invited to becode a developer at this time to integrate his work on ATAPI and ext2fs. After that, he kept working in the ATA and SCSI areas (he added the primary DMA support for PCI ATA controllers for example). On occasion, he also worked on other parts of the kernel (Ethernet controller drivers, of dual-endian support to FFS, for example).
When Xen 2.0 was out, he started working on domain0 support so that he could build virtual servers not relying on Linux. He added support for Xen 3.0 when it got out, and became NetBSD/xen portmaster at about this time. He also became a releng member in early 2007 to give some manpower to the team. He will give an overview of NetBSD/Xen and the new 'no-emulation' eltorito boot method.
Phil Nelson installed and ran the first web server serving www.netbsd.org. (the service has since been taken over by project servers.) He still does some minor admin work for the web server.
He runs the WWU build cluster and improved the scripts for the build clusters so they could build releases faster. He wrote a program called xarcmp that compares the contents of archives without having to extract them. It is part of the standard tool now used by both build clusters.
He wrote the initial versions of menuc, msgc and sysinst. sysinst is the current install system and it uses menuc and msgc. He will talk about LFS (Log-structured File System).
Julio M. Merino Vidal has been a member of The NetBSD Foundation and a developer since November 2002. He started contributing to the pkgsrc project with the main goal of getting GNOME 2 to work under NetBSD. He then made other contributions to the core NetBSD operating system, the most relevant of which are the tmpfs file system and the automated testing framework (ATF). He will discuss the features of tmpfs.
Antti Kantee worked on pkgsrc, device drivers and hardware support for various platforms, file systems, and assorted kernel and userland bits. He maintains the file(1) utility in the NetBSD source tree and NetHack in pkgsrc. He will describe the relationship between puffs, FUSE, and ReFUSE.
Alistair Crooks is the president of The NetBSD Foundation, Core Team member, and founder of pkgsrc.
He contributed to pkgsrc, user management software - user(8), ReFUSE (BSD-licensed re-implementation of FUSE), numerous filesystems based on ReFUSE, and iSCSI target and initiator. He will make clear which parts of the iSCSI protocol are included in this release and how much "hackathons" are helpful.
Reinoud Zandijk will add details about the implementation of the Universal Disk Format (UDF).
Iain Hibbert's major contribution has been the Bluetooth protocol stack and associated drivers and utilities. He talks about Bluetooth support in NetBSD.
Arnaud Lacombe has been a developer for about one year now. At the beginning, he joined the project to fix bugs found by the Coverity scan of NetBSD. He is interested in porting NetBSD to new embedded platforms. OpenMoko was a really interesting way to do it, but he never managed to get hardware, so this project is stalled for the moment.
Martin Fouts, Noud de Brouwer, and Arnaud discuss the current support for the hardware included in the iPhone and OpenMoko (Neo1973).
Alan Ritter will comment on the possibility of using the NDIS wrapper to port binary drivers among platforms.
Yamamoto Takashi will describe what agr(4) is.
Jan Schaumann used to work a lot on the web site. He also ported pkgsrc to IRIX and did bulk-builds there. He is a member of the communication-exec team and managed NetBSD's participation in the first two Google Summer of Code. He will discuss their experience with the Google Summer of Code project.
Release engineering, Sendmail, and kauth
What happened to the Release Engineering process for 4.0?
Liam J. Foy: Basically, the release engineering was started as planned. However, after the release engineering started a lot of changes were made which would be too time consuming to pull up (pull up from current to the NetBSD 4branch). Thus, we started the process again with all the changes merged.
You may remember a few NetBSD hackathons which took place. Well, these are what caused the large number of changes.
Why did you remove sendmail?
Christos Zoulas: Sendmail has been, is, and will be a security accident waiting to happen (unless it is rewritten from the ground up with security consciousness). Performing character pointer gymnastics in 50-100 line loops does not create any warm and fuzzy feelings for me. To top this off, most sendmail security issues are marked as confidential, and we are prevented from fixing or mentioning the problem until the ban is lifted. The last time this happened, we said "enough" and removed it altogether before the ban for that particular security issue was lifted
Would you like to present kauth, the kernel authorization framework, included in this release?
Elad Efrat: Kernel authorization is really something cool that I'm happy many people will get a chance to test in this release.
Basically, the story goes like this: the traditional Unix security model, the one where root is almighty and everyone else is not, is slowly beginning to show signs of age as demands for finer-grained security policies arise. The problem becomes more apparent when you look at the kernel code and see that, probably as time went by, people used various ways to check whether a certain operation is allowed to be performed or not—most are variants of checking whether the effective user-id is 0 or using the suser() function—effectively embedding the security model to the kernel, in a way that the question "Can this user open a raw socket?" really became "Is this user root?"
Kauth(9) was originally designed by Apple, and it provides a dispatching model that allows us to really ask if a certain operation can be performed by specifying it along with the relevant context. This abstracts the security model from the various kernel subsystems. The system is divided to "scopes" that collect actions of the same nature (right now, NetBSD has scopes like "process", "network","machdep", etc.), and each scope has a set of "actions" that define the operation (the context provided depends on the operation). When the kernel wants to ask if a certain operation is allowed, it calls the authorization wrapper for the relevant scope, providing it the action and context, and receives back a binary response of "yes" or "no".
This upper layer, like I said, effectively abstracts the way these decisions are made, allowing one to "plug" (almost) any security model one can think of.
These security models are implemented by attaching "listeners" (kauth(9) terminology for callback functions) that receive the action and context and make the decision. In the future it is hoped that these listeners will be able to run on an entirely different machine, thus allowing a centralized security policy for an entire network to be controlled from a single host.
Unfortunately, in NetBSD 4.0, although the security model abstraction is almost complete, there is still not full control over all operations. That is—in some places, the question is still "Is this user root?", so implementing different security models will not be complete. However, it will provide an indication of how well kauth(9) works and how users react to this new ability to develop custom security models (for example, classic uses like restricting raw networking, or binding to certain low ports, to a few users). I'm hoping to be surprised by our user base's creativity.
I highly recommend spending a few minutes reading kauth(9) and secmodel(9) in the current man pages. There's a little taste of what's in the plans for NetBSD 5.0—like credential inheritance control—with the ultimate goal being to provide system administrators with security policy control that requires either zero or very little effort.
PaX, fileassoc, and Veriexec
What security features have you added to mprotect(2) from PaX?
Elad Efrat: First, the PaX project, for those who aren't familiar with it, is responsible for all the modern exploit mitigation technologies. It's where stuff like W^X and ASLR (Address Space Layout Randomization) were born, along with many, many other cool features. One of them, which I initially ported to NetBSD, is PaX MPROTECT.
PaX MPROTECT can be thought of as "strict W^X". If in a normal system, a program starts with no memory pages that are both writable and executable, but those pages can still be created using mprotect(2), systems or programs that run with PaX MPROTECT are "immune" to attacks where protection on pages is modified, often by trashing arguments to mprotect(2).
paxctl(8) is a tool that enables PaX MPROTECT on a per-program basis, if you don't want to enable it globally. I'm afraid that not much testing was done on various architectures using this feature, so users should first experiment with it...
How does the new fileassoc KPI (Kernel Programming Interface) work?
Elad Efrat: I'll begin by describing what fileassoc(9) is.
Some (newer) filesystems allow on to attach metadata to files by using extended attributes (in addition to the usual ones—permissions, timestamps, etc.) and such. A common use for them may be, for example, to store ACLs (Access Control List) for the file that allow finer granularity on access control.
The problem is that these extended attributes are filesystem-dependent and either don't exist on most filesystems or are interfaced differently. This is where fileassoc(9) comes into play: it allows you to attach metadata to files in a filesystem-independent way while storing the information in fast-access kernel tables. The advantage is that it really is filesystem-independent, so there's the potential of adding features and such to filesystems that lack them (for example, again, ACLs). The disadvantages are that the metadata is not really "attached" to the file in any way, and must be loaded in some form—from a database file or such—to the kernel. So that the dependency becomes on the OS itself, but that's cool, because it's aimed for NetBSDsystems.
The KPI works by storing data in hash tables. Metadata is identified by a key (a C string) and is a stream of bytes. A kernel subsystem can then use the fileassoc(9) KPI to either attach, query, or remove metadata from files, by specifying the key.
The fileassoc(9) KPI is implemented in the VFS (Virtual File System) layer, identifying each file using its "file handle" (these are supposed to be unique). So pretty much any filesystem is supported. An example (and the only, at the moment) consumer of fileassoc(9) is Veriexec, where a database file holds information that is parsed by a userland tool, then fed to fileassoc(9) using a special device.
Plans to write a generic interface for communicating with fileassoc(9) have come up in the past, but since we're not yet convinced of the necessity, these are just plans at the moment. Developers interested in feeding data to fileassoc(9) should also implement their own special device stub—at least for now.
Is there any news about Veriexec?
Elad Efrat: As with NetBSD 4.0, I feel Veriexec has gotten to a point where it should probably be used by most of our users.
Veriexec is NetBSD's integrity subsystem, which, in short, can guarantee the integrity of the programs, configuration files, etc., on the system.
In addition to the tons of improvements in performance and stability, and the many features added, NetBSD 4.0 introduces 'veriexecgen', which is a tool written by Matt Fleming (mjf@). Veriexecgen tremendously lowers the bar for fingerprint database generation to a point where it's possible to run a single command with no arguments after installation to have Veriexec set-up appropriately. I strongly recommend reading the veriexecgen(8) main page and giving it a try.
What is veriexecgen and how does it work?
Matt Fleming: veriexecgen is a program that runs against directories and generates a set of fingerprints for use with veriexec. This fingerprint database is usually stored in /etc/signatures.
XFree86, pkgsrc, proplib, and Xen
Do you still include Linux binary compatibility?
Nicolas Joly: I'm currently working on improving Linux binary compatibility on AMD64 for both 32- and 64-bit applications.
The amd64 compat linux is the first one that include NPTL (NativePOSIX Thread Library) emulation found on Linux 2.6 kernel. Likewise, compat linux32 is pretty new, and even if mostly identical to i386linux emulation needs to be modified.
For NetBSD 4.0, kernel support for compat linux is not enabled bydefault as this is not as stable as other ports. For this release, do not expect to run complicated linux applications yet, but basic ones should work. In the mean time, -current has made some progresses...
Which X Window System is included in NetBSD 4.0?
Matthias Scheler: XFree86 4.5.0.
The XFree86 Project has unfortunately lost a lot of momentum. NetBSD is currently in the process of switching to the X.org X11 distribution. We initially tried integrating the monolithic X.org distribution into our xsrc source tree. But it never reached a state where it worked on all platforms and supported cross builds via build.sh.
After the X.org project changed to a modular distribution, it was obvious that pkgsrc is the best way to integrate X11 in the future. There is ongoing work mostly by Joerg Sonnenberger to make modular X11 in pkgsrc cross-buildable. When that work has been finished and properly integrated with the system installation, the XFree86 distribution in xsrc will probably be retired.
What's new in pkgsrc?
Joerg Sonnenberger: I'm only aware of NetBSD 4 being a requirement to use the cross-compiling support. The updates to pkg_install will be in NetBSD4.1, so that is ruled out.
The cross-compiling support allows a small subset of pkgsrc to be built for any architecture running NetBSD using the output of build.sh. Currently this subset includes modular Xorg and a few applications like Xpdf as proof of concept. This is intended to replace the aging XFree86in xsrc, but needs some more polishing work on the pkgsrc side.
What is proplib(3) and what can we use it for?
Jason Thorpe: proplib is a library for manipulating property lists. Property lists are collections of properties, typically stored in a dictionary. A dictionary is an associative array, essentially key-valuepairs. The keys are strings, and the values are strings, numbers, opaque data, booleans, arrays, dictionaries, ...
This is handy in a variety of applications. For example, communicating structured data between userspace and the kernel, describing properties of devices, etc.
What can we do with NetBSD/Xen?
Manuel Bouyer: Well, we can do a lot of things, from merging different physical boxes to a single one (if you need different OSes for different tasks, for example), to easily create test systems (it's very convenient for kernel developement: a guest boots much faster than a regular PC). With the virtualization features of recent X86 CPUs (which are supported by Xen3 with NetBSD as dom0), you can also boot plain i386 OSes. This can be used to test install media, for example, or to load systems that can't be paravirtualized (e.g., Windows).
Filesystems
What is the status of LFS (Log-structured File System)?
Phil Nelson: LFS is in use on about 1/2 of the 23 nodes of the WWU build cluster. Granted, it is running an older 4.0-Beta kernel, but they have been running just fine for quite a while. Currently all the machines in the WWU build cluster have been up 78+ days and continuously building. They are all i386 machines.
Would you like to describe the features of tmpfs?
Julio M. Merino Vidal: tmpfs was born as a replacement for MFS—the memory file system included in all BSDs as far as I know—and as such it shares many functionality with it. It basically is an efficient, memory-based file system which means that it uses part of the system's virtual memory to store files in a way that is more space- and speed-efficient than MFS. It supports all of the features expected from a Unix-style filesystem, including hard links, symbolic links, devices, permissions, file flags and NFS exportability, but it currently does not support sparse files.
As opposed to MFS, tmpfs file systems can grow and shrink automatically depending on the available memory if they are configured as such; their memory consumption can also be upper-limited. And even though the end user cannot directly notice it, tmpfs's code is much simpler than MFS's one, which means that it is easier to audit, test and optimize.
As you just mentioned, tmpfs is a bit different than MFS, but why did MFS need a replacement in the first place?
Julio M. Merino Vidal: The main problem of MFS is that it performs poorly, uses too much memory, and the used memory cannot be reclaimed unless the file system is unmounted. To understand why this happens, we need to see how MFS works, and to see that, we first need to outline how FFS is designed.
FFS (or UFS), the traditional BSD file system, is conceptually split in two different layers. the upper layer handles all the logic from the file system, including the management of directories, the common file operations, the routines to map data to blocks, etc. The lower layer lays out the data on disk, allocating blocks and inodes as needed, etc. This seems like a very nice design, but in the end, it imposes restrictions on the way the lower layer operates (or at least that's the impression I gathered; I'm not too familiar on FFS's code).
MFS is just a replacement of FFS's lower layer to operate on virtual memory. The approach it takes is dead simple: it allocates a contiguous block of memory and treats that as if it were a disk, organizing the contents of such memory region as disk blocks. As you can expect, such approach has a lot of overhead because the file system uses an incorrect abstraction—disk blocks and cylinder groups—to manage memory.
tmpfs, on the other hand, is a much simpler filesystem which takes advantage of the fact that it is always memory-backed. It uses traditional structures, linked lists, arrays, memory pools and other sorts of data abstractions to represent the contents of the file system. These abstractions are easier to deal with and hence require less resources to manage.
It is also interesting to mention that tmpfs grew a custom regression testing suite that was also well received by developers. But the way it was created was suboptimal, which made me start the ATF (Automated Testing Framework) project this year. We'll see a better testing suite in NetBSD 5.0, or at least I hope so! Oh, and by the way, FreeBSD also imported tmpfs in its source tree.
What is the relationship between puffs, FUSE and ReFUSE?
Antti Kantee: puffs stands for Pass-to-Userspace Framework File System. It is an interface and a framework implementation for userspace filesystems on NetBSD. The userspace filesystem interface itself is heavily influenced by the kernel virtual filesystem interface. puffs can be thought to consist of two parts: the mechanism for transporting file system requests from the kernel to userspace and a userspace library, libpuffs, for interfacing with the kernel and writing the userspace filesystem implementation.
FUSE, or Filesystem in Userspace, is another API for implementing userspace filesystems. It is native to Linux, but has been widely ported to other operating systems such as FreeBSD and OpenSolaris. FUSE is considered the standard interface for implementing userspace filesystems with numerous filesystems readily available.
ReFUSE is the implementation of the FUSE API for NetBSD. It is implemented on top of puffs and completely done in userspace. Other non-Linux operating systems implement FUSE compatibility in the kernel, but we believe the ReFUSE approach is the right way to address the issue: export the kernel filesystem interface in the most natural way possible (puffs) and implement compatibility in userspace (ReFUSE).
A very important difference to note is API stability. Since the FUSE API is well-established, it is stable, and filesystems written for it will most likely continue to function for a long time. The puffs API is a "low-level" API which follows the NetBSD kernel virtual filesystem API fairly closely. As the virtual filesystem API on NetBSD evolves, so will the puffs API (and vice versa, if things go as planned). A benefit from this evolutionary incompatibility is the ability to use all the features provided by the kernel API at the best possible performance.
What is their status in NetBSD 4.0?
Antti Kantee: puffs is still under heavy development. The version of puffs found in 4.0 is a snapshot of what was in the tree at the time 4.0 was branched. This means that ReFUSE and therefore FUSE support is unfortunately not present. The most useful application for users is likely to be ssshfs, simple sshfs, which can be found in source form from the tree under src/share/examples/puffs/ssshfs. It implements sshfs functionality. This implementation was superceded by mount_psshfs(8) for NetBSD 5.0.
Technically there is no reason why support for FUSE could not be added to the NetBSD 4 branch. However, it requires backporting ReFUSE and some puffs features from the development branch and therefore a person with time and motivation to do the work, and, above all, test and support it.
Friday, February 8, 2008
Autosave script for Keynote
Apple Tech Forums
As many of you may already realize there is no autosave feature that is native to Keynote for the Mac; so after searching around for a few I found this little script that should do the trick for you. Good Luck!!!
Hunter
Tired to press cmd + S ?
to autosave iWork's documents
paste this script in the Script Editor
menu > File > Save As…
uncheck screen at startup
check stay in background
save this script as an application
Install it as one of the applications launched at boot_time from the Accounts PreferencePane (tab Login Items)
Every ten minutes, if Pages or Keynote is in use, the open documents (already saved once) would be saved.
if Numbers is in use, the frontmost document will be saved.
If I made no mistake, if the document was never saved the "Save" dialog will ask you to save it. (OK, it works)
--SCRIPT autosave4iWork
property minutesBetweenSaves : 10
on idle
my auto4PK("Pages")
my auto4PK("Keynote")
my auto4Numbers()
return minutesBetweenSaves * 60
end idle
on auto4PK(theApp)
tell application "System Events" to set theAppIsRunning to (name of processes) contains "theApp"
if theAppIsRunning = true then
tell application theApp
repeat with aDoc in every document
if path of aDoc exists then
if modified of aDoc then save aDoc
end if
end repeat
end tell
end if
end auto4PK
on auto4Numbers()
set theApp to "Numbers"
tell application "System Events" to set theAppIsRunning to (name of processes) contains theApp
if theAppIsRunning = true then
tell application theApp to activate
tell application "System Events" to tell application process theApp to keystroke "s" using {command down}
end if
end auto4Numbers
--SCRIPT
As many of you may already realize there is no autosave feature that is native to Keynote for the Mac; so after searching around for a few I found this little script that should do the trick for you. Good Luck!!!
Hunter
Tired to press cmd + S ?
to autosave iWork's documents
paste this script in the Script Editor
menu > File > Save As…
uncheck screen at startup
check stay in background
save this script as an application
Install it as one of the applications launched at boot_time from the Accounts PreferencePane (tab Login Items)
Every ten minutes, if Pages or Keynote is in use, the open documents (already saved once) would be saved.
if Numbers is in use, the frontmost document will be saved.
If I made no mistake, if the document was never saved the "Save" dialog will ask you to save it. (OK, it works)
--SCRIPT autosave4iWork
property minutesBetweenSaves : 10
on idle
my auto4PK("Pages")
my auto4PK("Keynote")
my auto4Numbers()
return minutesBetweenSaves * 60
end idle
on auto4PK(theApp)
tell application "System Events" to set theAppIsRunning to (name of processes) contains "theApp"
if theAppIsRunning = true then
tell application theApp
repeat with aDoc in every document
if path of aDoc exists then
if modified of aDoc then save aDoc
end if
end repeat
end tell
end if
end auto4PK
on auto4Numbers()
set theApp to "Numbers"
tell application "System Events" to set theAppIsRunning to (name of processes) contains theApp
if theAppIsRunning = true then
tell application theApp to activate
tell application "System Events" to tell application process theApp to keystroke "s" using {command down}
end if
end auto4Numbers
--SCRIPT
Thursday, February 7, 2008
Three photo mosaic apps compared
By Nathan Willis on February 04, 2008 (4:00:00 PM)
Photo mosaics are recreations of one large image composed of tiny tiles of other smaller images. They can be a fun project and make good use of the hundreds of less-than-extraordinary photos on your hard drive. We compared three easy-to-use Linux-based utilities for generating photo mosaics -- Pixelize, Metapixel, and Imosaic -- on speed, quality, and other factors.
Pixelize
Pixelize is the oldest of the three apps; it uses GTK1 for its interface, which may make it stand out from your other desktop applications. You won't be able to use your favorite file selection widget, either, which can be frustrating -- the GTK1 file selector does not remember the last directory you visited, and it does not provide thumbnail previews of image files.
Pixelize in action
You can download source code from the project's Web page, but check your Linux distribution's package management system first -- many distros include Pixelize. The latest version is 0.9.2, and does not have any unusual package prerequisites.
The package includes two programs: the GUI front end pixelize, and the command-line back end make_db. You must run make_db at least once before you run pixelize in order to populate the database of source images from which the app will tile your mosaic. Just run make_db somedirectory/*, and make_db will index all of the image files within somedirectory. The utility cannot descend into other directories recursively, so you will have to execute make_db multiple times if you store your images in multiple locations.
Once you are ready, launch the GUI with pixelize &. Up will pop an empty window; choose File -> Open to select the picture from which you want to create a mosaic.
Choose Options -> Options to adjust the program's parameters. You can specify either the number of images to use in each row and column of the mosaic, or the size to make each mosaic tile -- changing either factor automatically adjusts the other. The only other option to worry about is the proximity of duplicates. If you have a very small set of source images, Pixelize will have to use each several times; this parameter allows you to space them further apart to make them less visible.
Once you are ready, choose Options -> Render, and Pixelize will generate your mosaic, drawing it on top of the original image in the main window. If you like what you see, choose File -> Save.
Metapixel
Metapixel is a command-line-only tool, but it offers more flexibility in mosaic creation than does Pixelize. The latest release is version 1.0.2, which you can download from the project's site. Source code as well as Fedora RPMs are available. But as with Pixelize, many modern Linux distros ship Metapixel, so see if you can use an official package if you are not interested in compiling your own binary. Metapixel requires Perl for its image preparation step.
Before you can create a mosaic, you must run metapixel-prepare to prepare a batch of source images. The complete syntax is metapixel-prepare --width=n --height=m --recurse sourcedirectory destinationdirectory. This will scan through sourcedirectory (recursively), and generate n-by-m tile images for each image it finds, saving them to destinationdirectory. The --recurse flag is optional, as are --width and --height; without them metapixel-prepare will create 128x128 tiles.
You can create a basic mosaic with metapixel --library=destinationdirectory --metapixel inputimage.jpg outputimage.jpg. The --library flag should designate a directory of tiles preprocessed by metapixel-prepare; the ability to maintain multiple such directories allows you more flexibility in building your mosaics. Metapixel can generate JPEG or PNG output, depending on the output file name you provide.
Metapixel has several optional flags you can use to tweak the output. You can scale the output file in relation to the input with --scale=x, or assign different relative weights to the different color channels in the pattern-matching step. Collage mode (invoked with --collage) generates a mosaic where the tiles overlap each other, rather than being in strict rows and columns. There are even two different algorithms for determining which tile best matches which part of the mosaic, so you can experiment and decide what produces the best results for your input and tile collection.
Imosaic
IMosaic in action
Imosaic is the only closed source app in this mix. It is a .Net program that the author has also packaged for Linux and Mac OS X to run using Mono. You can download the latest package from imosaic.net. Unpack the archive anywhere on your system and launch the app with mono IMosaic.exe &.
I had trouble getting the software to run using Ubuntu 7.10's Mono; even with all of the available mono* packages, Imosaic refused to run. I was able to get it working by downloading a vanilla Mono installer from the Mono Web site; luckily you can have more than one copy of Mono installed at any time.
Once you launch Imosaic, the first step to creating a mosaic is defining a collection of tile images. Choose Edit -> Images Collection from the menu, and in the Image Collection Editor, add as many files and directories as you need. To use your collection, you must explicitly choose Task -> Process images to prepare the tiles, then explicitly save the collection via File -> Save Collection. When you have a collection defined, you add it to the mosaic creation process under Imosaic's Image's Collections [sic] tab.
You select the image you want to transform into a mosaic with File -> Open Source Image. By default Imosaic displays grid lines over the image to show you where the mosaic tiles will fall, which is helpful. You can resize the tile dimensions in the Settings tab, as well as adjust factors like JPEG output quality and distance between duplicate tiles. When you are ready to convert, click the Process button and wait. The Statistics tab provides help info on the mosaic creation process.
The Linux and OS X port of Imosaic is recent, and that is reflected in UI quirks that take some getting used to. For instance, Windows naming conventions are used throughout the interface, such as My Computer and Personal -- the latter of which refers to your home folder. I also encountered a lot of sudden crashes when trying to build image collections; after every three or four "Add directory" steps the app would blink out of existence, erasing all changes I had made to my in-progress collection. Until the developer fixes the problem, consider yourself warned and save those collections often.
Triptych
Side by side, it is hard to pick a best program from among the three apps. Metapixel offers far more mosaic creation options, but its lack of a GUI means you have to use trial and error to get a good feel for what those options do. I like how simple Pixelize is and how you can see your mosaic being generated as it happens on screen, but GTK1 is ugly and almost unusable nowadays.
Imosaic bests both of the others in a few important areas. It is the only program that integrates preprocessing an image collection (or even informs you in the interface that you need to preprocess a collection), and is the only one that allows you to cancel mosaic creation mid-process. But the Linux build is so buggy for now that it is hard to build a large enough collection to produce good output, and several of its key features (e.g. the Sequences tab) remain undocumented and are thus of unknown value.
Left to right: output from Pixelize, Metapixel, and IMosaic
As far as output quality goes, I was happiest with Metapixel's results, using the --metric=wavelet and --search=global options. Pixelize tended to produce desaturated, nearly monochrome mosaics -- and I have thousands of color images, so the problem is likely with the program's algorithm. Imosaic produced respectable results, but the frequent crashes made it hard to build a large image collection with the variety needed to generate large mosaics.
A good GUI for Metapixel and some bug fixes for Imosaic would put them on approximately equal footing.
Finally, if you take an interest in photo mosaics, you might be interested to learn of the patent situation. Several patents on creating photo mosaics were granted to US-based Runaway Technology, beginning in 2000, despite apparent evidence of prior art dating back at least as far as 1993. It is unclear if the company has successfully defended its patents, but if your dreams include monetizing your mosaics, the inventions claimed in the patent are something you should look into.
Photo mosaics are recreations of one large image composed of tiny tiles of other smaller images. They can be a fun project and make good use of the hundreds of less-than-extraordinary photos on your hard drive. We compared three easy-to-use Linux-based utilities for generating photo mosaics -- Pixelize, Metapixel, and Imosaic -- on speed, quality, and other factors.
Pixelize
Pixelize is the oldest of the three apps; it uses GTK1 for its interface, which may make it stand out from your other desktop applications. You won't be able to use your favorite file selection widget, either, which can be frustrating -- the GTK1 file selector does not remember the last directory you visited, and it does not provide thumbnail previews of image files.
Pixelize in action
You can download source code from the project's Web page, but check your Linux distribution's package management system first -- many distros include Pixelize. The latest version is 0.9.2, and does not have any unusual package prerequisites.
The package includes two programs: the GUI front end pixelize, and the command-line back end make_db. You must run make_db at least once before you run pixelize in order to populate the database of source images from which the app will tile your mosaic. Just run make_db somedirectory/*, and make_db will index all of the image files within somedirectory. The utility cannot descend into other directories recursively, so you will have to execute make_db multiple times if you store your images in multiple locations.
Once you are ready, launch the GUI with pixelize &. Up will pop an empty window; choose File -> Open to select the picture from which you want to create a mosaic.
Choose Options -> Options to adjust the program's parameters. You can specify either the number of images to use in each row and column of the mosaic, or the size to make each mosaic tile -- changing either factor automatically adjusts the other. The only other option to worry about is the proximity of duplicates. If you have a very small set of source images, Pixelize will have to use each several times; this parameter allows you to space them further apart to make them less visible.
Once you are ready, choose Options -> Render, and Pixelize will generate your mosaic, drawing it on top of the original image in the main window. If you like what you see, choose File -> Save.
Metapixel
Metapixel is a command-line-only tool, but it offers more flexibility in mosaic creation than does Pixelize. The latest release is version 1.0.2, which you can download from the project's site. Source code as well as Fedora RPMs are available. But as with Pixelize, many modern Linux distros ship Metapixel, so see if you can use an official package if you are not interested in compiling your own binary. Metapixel requires Perl for its image preparation step.
Before you can create a mosaic, you must run metapixel-prepare to prepare a batch of source images. The complete syntax is metapixel-prepare --width=n --height=m --recurse sourcedirectory destinationdirectory. This will scan through sourcedirectory (recursively), and generate n-by-m tile images for each image it finds, saving them to destinationdirectory. The --recurse flag is optional, as are --width and --height; without them metapixel-prepare will create 128x128 tiles.
You can create a basic mosaic with metapixel --library=destinationdirectory --metapixel inputimage.jpg outputimage.jpg. The --library flag should designate a directory of tiles preprocessed by metapixel-prepare; the ability to maintain multiple such directories allows you more flexibility in building your mosaics. Metapixel can generate JPEG or PNG output, depending on the output file name you provide.
Metapixel has several optional flags you can use to tweak the output. You can scale the output file in relation to the input with --scale=x, or assign different relative weights to the different color channels in the pattern-matching step. Collage mode (invoked with --collage) generates a mosaic where the tiles overlap each other, rather than being in strict rows and columns. There are even two different algorithms for determining which tile best matches which part of the mosaic, so you can experiment and decide what produces the best results for your input and tile collection.
Imosaic
IMosaic in action
Imosaic is the only closed source app in this mix. It is a .Net program that the author has also packaged for Linux and Mac OS X to run using Mono. You can download the latest package from imosaic.net. Unpack the archive anywhere on your system and launch the app with mono IMosaic.exe &.
I had trouble getting the software to run using Ubuntu 7.10's Mono; even with all of the available mono* packages, Imosaic refused to run. I was able to get it working by downloading a vanilla Mono installer from the Mono Web site; luckily you can have more than one copy of Mono installed at any time.
Once you launch Imosaic, the first step to creating a mosaic is defining a collection of tile images. Choose Edit -> Images Collection from the menu, and in the Image Collection Editor, add as many files and directories as you need. To use your collection, you must explicitly choose Task -> Process images to prepare the tiles, then explicitly save the collection via File -> Save Collection. When you have a collection defined, you add it to the mosaic creation process under Imosaic's Image's Collections [sic] tab.
You select the image you want to transform into a mosaic with File -> Open Source Image. By default Imosaic displays grid lines over the image to show you where the mosaic tiles will fall, which is helpful. You can resize the tile dimensions in the Settings tab, as well as adjust factors like JPEG output quality and distance between duplicate tiles. When you are ready to convert, click the Process button and wait. The Statistics tab provides help info on the mosaic creation process.
The Linux and OS X port of Imosaic is recent, and that is reflected in UI quirks that take some getting used to. For instance, Windows naming conventions are used throughout the interface, such as My Computer and Personal -- the latter of which refers to your home folder. I also encountered a lot of sudden crashes when trying to build image collections; after every three or four "Add directory" steps the app would blink out of existence, erasing all changes I had made to my in-progress collection. Until the developer fixes the problem, consider yourself warned and save those collections often.
Triptych
Side by side, it is hard to pick a best program from among the three apps. Metapixel offers far more mosaic creation options, but its lack of a GUI means you have to use trial and error to get a good feel for what those options do. I like how simple Pixelize is and how you can see your mosaic being generated as it happens on screen, but GTK1 is ugly and almost unusable nowadays.
Imosaic bests both of the others in a few important areas. It is the only program that integrates preprocessing an image collection (or even informs you in the interface that you need to preprocess a collection), and is the only one that allows you to cancel mosaic creation mid-process. But the Linux build is so buggy for now that it is hard to build a large enough collection to produce good output, and several of its key features (e.g. the Sequences tab) remain undocumented and are thus of unknown value.
Left to right: output from Pixelize, Metapixel, and IMosaic
As far as output quality goes, I was happiest with Metapixel's results, using the --metric=wavelet and --search=global options. Pixelize tended to produce desaturated, nearly monochrome mosaics -- and I have thousands of color images, so the problem is likely with the program's algorithm. Imosaic produced respectable results, but the frequent crashes made it hard to build a large image collection with the variety needed to generate large mosaics.
A good GUI for Metapixel and some bug fixes for Imosaic would put them on approximately equal footing.
Finally, if you take an interest in photo mosaics, you might be interested to learn of the patent situation. Several patents on creating photo mosaics were granted to US-based Runaway Technology, beginning in 2000, despite apparent evidence of prior art dating back at least as far as 1993. It is unclear if the company has successfully defended its patents, but if your dreams include monetizing your mosaics, the inventions claimed in the patent are something you should look into.
Sunday, February 3, 2008
Efficient rsyncrypto hides remote sync data
By Ben Martin on February 01, 2008 (9:00:00 AM)
The rsync utility is smart enough to send only enough bytes of a changed file to a remote system to enable the remote file to become identical to the local file. When that information is sensitive, using rsync over SSH protects files while in transit.To protect the files when they are on the server you might first encrypt them with GPG. But the manner in which GPG encrypts slightly changed files foils rsync's efficiency.rsyncrypto allows you to encrypt your files while still allowing you to leverage the speed of rsync.
One of the aims of encryption is to try to make any change to the unencrypted file completely modify how the encrypted file appears. This means that somebody who has access to a series of encrypted files gets little information about how the little changes you might make to the unencrypted file are affecting the encrypted file over time. The downside is that such security causes the encryption program to modify most of the encrypted file. If you then use rsync to copy such a file to the remote server, it will have to send almost the complete file to the remote server each time.
The goal of rsyncrypto is to encrypt files in such a manner that only a slight and controlled amount of security is sacrificed in order to make rsync able to send the encrypted files much quicker. It aims to leak no more than 20 bits of aggregated information per 8KB of plaintext file.
For an example of an information leak, suppose you have an XML file and you use rsyncrypto to copy the file to a remote host. Then you change a single XML attribute and use rsyncrypto to copy the updates across. Now suppose an attacker captured the encrypted versions in transit, and thus has copies of both the encrypted file before the change and after the change. The first thing they learn is that only the first 8KB of the file changed, because that is all that was sent the second time. If they can speculate what sort of file the unencrypted file was (for example, an XML file) then they can try to use that guess in an attempt to recover information.
Rsyncrypto encrypts parts of the file independently, thus keeping any changes you make to a single block of the file local to that block in the encrypted version. If you're protecting a collection of personal files from a possible remote system compromise, such a tradeoff in security might be acceptable. On the other hand, if you cannot allow any information leaks, then you'll have to accept that the whole encrypted file will change radically each time you change the unencrypted file. If that's the case, using rsync on GnuPG-encrypted files might suit your needs.
On a Fedora 8 machine, you have to download both rsyncrypto and the dependency argtable2 and install them using the standard ./configure; make; make install combination, starting with argtable2.
rsyncrypto is designed to be used as a presync option to rsync. That is, you first use rsyncrypto on the plain unencrypted files to obtain an encrypted directory tree, and you then run rsync to send that encrypted tree to a remote system. The following command syntax shows a template for directory encryption and decryption:
# to encrypt rsyncrypto -r srcdir /tmp/encrypted srcdir.keys mykey.crt # to decrypt rsyncrypto -d -r /tmp/encrypted srcdir srcdir.keys mykey.crt
The keys and certificates referenced in these commands are generated by OpenSSL, as we'll see in a moment. In the commands, the srcdir is encrypted and sent to /tmp/encrypted with the individual keys used to encrypt each file in srcdir saved into srcdir.keys. The mykey.crt is a certificate that is used to protect all the keys in srcdir.keys. If you still have all the keys, you can use your certificate in the decryption operation to obtain the plaintext files again. If you lose srcdir.keys, all is not lost, but you must use the private key for mykey.crt to regain the encrypted keys that are also stored in /tmp/encrypted.
The following is a full one-way sync to a remote server using both rsyncrypto and rsync to obtain an encrypted backup on a remote machine. The example first generates a master key and certificate using OpenSSL, then makes an encrypted backup of ~/foodir onto the remote machine v8tsrv:
$ mkdir ~/rsyncrypto-keys $ cd ~/rsyncrypto-keys $ openssl req -nodes -newkey rsa:1536 -x509 -keyout rckey.key -out rckey.crt $ cd ~ $ mkdir foodir $ date >foodir/df1.txt $ date >foodir/df2.txt $ rsyncrypto -r foodir /tmp/encrypted foodir.keys ~/rsyncrypto-keys/rckey.crt $ rsync -av /tmp/encrypted ben@v8tsrv:~
In order to test the speed gain of using rsyncrypto as opposed to using other encryption with rsync, I used /dev/urandom to create a file of random bytes, encrypted it with both rsyncrypto and GnuPG, and rsynced both of these to a remote system using rsync. I then modified the plaintext file, encrypted the file again, and synced the encrypted file with the remote system. In this case, I modified 6KB of data at an offset of 17KB into the file using dd and left all the other data intact. The final rsync commands show that the rsyncrypto-encrypted tree only needed to send 58,102 bytes, whereas the GnuPG-encrypted file required the entire file to be sent to the remote system:
$ cd ~/foodir $ rm -f * $ dd if=/dev/urandom of=testfile.random bs=1024 count=500 512000 bytes (512 kB) copied, 0.088045 s, 5.8 MB/s $ cd ~ $ rsyncrypto -r foodir foodir.rcrypto foodir.keys ~/rsyncrypto-keys/rckey.crt $ ls -l foodir.rcrypto/ -rw-r--r-- ... 502K 2008-01-08 19:59 testfile.random $ mkdir foodir.gpg $ gpg --gen-key ... $ mkdir foodir.gpg $ gpg --output foodir.gpg/testfile.random.gpg -e foodir/testfile.random $ ls -l foodir.gpg -rw-r--r-- 1 ... 501K 2008-01-08 20:07 testfile.random.gpg $ rsync -av foodir.rcrypto ben@v8tsrv:~ sent 513356 bytes received 48 bytes 342269.33 bytes/sec $ rsync -av foodir.gpg ben@v8tsrv:~ sent 513026 bytes received 48 bytes 342049.33 bytes/sec # # modify the input file starting at 17KB into the file for 6KB # $ dd if=/dev/urandom of=~/foodir/testfile.random bs=1024 count=6 seek=17 conv=notrunc $ ls -l testfile.random -rw-r--r-- 1 ... 500K 2008-01-08 20:17 testfile.random $ rsyncrypto -r foodir foodir.rcrypto foodir.keys ~/rsyncrypto-keys/rckey.crt $ gpg --output foodir.gpg/testfile.random.gpg -e foodir/testfile.random # # See how much gets sent # $ rsync -av foodir.rcrypto ben@v8tsrv:~ sent 58102 bytes received 4368 bytes 124940.00 bytes/sec $ rsync -av foodir.gpg ben@v8tsrv:~ sent 513024 bytes received 4368 bytes 1034784.00 bytes/sec
Using rsyncrypto with rsync, you can protect the files that you send to a remote system while allowing modified files to be sent in a bandwidth-efficient manner. There is a slight loss of security using rsyncrypto, because changes to the unencrypted file do not propagate throughout the entire encrypted file. When this security trade-off is acceptable, you can get much quicker bandwidth-friendly network syncs and still achieve good encryption on the files stored on the remote server.
If you wish to hide file names as well as their content on the remote server, you can use the --name-encrypt=map option to rsyncrypto, which stores a mapping from the original file name to a garbled random file name in a mapping file, and outputs files using only their garbled random file name in the encrypted directory tree.
Ben Martin has been working on filesystems for more than 10 years. He completed his Ph.D. and now offers consulting services focused on libferris, filesystems, and search solutions.
The rsync utility is smart enough to send only enough bytes of a changed file to a remote system to enable the remote file to become identical to the local file. When that information is sensitive, using rsync over SSH protects files while in transit.To protect the files when they are on the server you might first encrypt them with GPG. But the manner in which GPG encrypts slightly changed files foils rsync's efficiency.rsyncrypto allows you to encrypt your files while still allowing you to leverage the speed of rsync.
One of the aims of encryption is to try to make any change to the unencrypted file completely modify how the encrypted file appears. This means that somebody who has access to a series of encrypted files gets little information about how the little changes you might make to the unencrypted file are affecting the encrypted file over time. The downside is that such security causes the encryption program to modify most of the encrypted file. If you then use rsync to copy such a file to the remote server, it will have to send almost the complete file to the remote server each time.
The goal of rsyncrypto is to encrypt files in such a manner that only a slight and controlled amount of security is sacrificed in order to make rsync able to send the encrypted files much quicker. It aims to leak no more than 20 bits of aggregated information per 8KB of plaintext file.
For an example of an information leak, suppose you have an XML file and you use rsyncrypto to copy the file to a remote host. Then you change a single XML attribute and use rsyncrypto to copy the updates across. Now suppose an attacker captured the encrypted versions in transit, and thus has copies of both the encrypted file before the change and after the change. The first thing they learn is that only the first 8KB of the file changed, because that is all that was sent the second time. If they can speculate what sort of file the unencrypted file was (for example, an XML file) then they can try to use that guess in an attempt to recover information.
Rsyncrypto encrypts parts of the file independently, thus keeping any changes you make to a single block of the file local to that block in the encrypted version. If you're protecting a collection of personal files from a possible remote system compromise, such a tradeoff in security might be acceptable. On the other hand, if you cannot allow any information leaks, then you'll have to accept that the whole encrypted file will change radically each time you change the unencrypted file. If that's the case, using rsync on GnuPG-encrypted files might suit your needs.
On a Fedora 8 machine, you have to download both rsyncrypto and the dependency argtable2 and install them using the standard ./configure; make; make install combination, starting with argtable2.
rsyncrypto is designed to be used as a presync option to rsync. That is, you first use rsyncrypto on the plain unencrypted files to obtain an encrypted directory tree, and you then run rsync to send that encrypted tree to a remote system. The following command syntax shows a template for directory encryption and decryption:
# to encrypt rsyncrypto -r srcdir /tmp/encrypted srcdir.keys mykey.crt # to decrypt rsyncrypto -d -r /tmp/encrypted srcdir srcdir.keys mykey.crt
The keys and certificates referenced in these commands are generated by OpenSSL, as we'll see in a moment. In the commands, the srcdir is encrypted and sent to /tmp/encrypted with the individual keys used to encrypt each file in srcdir saved into srcdir.keys. The mykey.crt is a certificate that is used to protect all the keys in srcdir.keys. If you still have all the keys, you can use your certificate in the decryption operation to obtain the plaintext files again. If you lose srcdir.keys, all is not lost, but you must use the private key for mykey.crt to regain the encrypted keys that are also stored in /tmp/encrypted.
The following is a full one-way sync to a remote server using both rsyncrypto and rsync to obtain an encrypted backup on a remote machine. The example first generates a master key and certificate using OpenSSL, then makes an encrypted backup of ~/foodir onto the remote machine v8tsrv:
$ mkdir ~/rsyncrypto-keys $ cd ~/rsyncrypto-keys $ openssl req -nodes -newkey rsa:1536 -x509 -keyout rckey.key -out rckey.crt $ cd ~ $ mkdir foodir $ date >foodir/df1.txt $ date >foodir/df2.txt $ rsyncrypto -r foodir /tmp/encrypted foodir.keys ~/rsyncrypto-keys/rckey.crt $ rsync -av /tmp/encrypted ben@v8tsrv:~
In order to test the speed gain of using rsyncrypto as opposed to using other encryption with rsync, I used /dev/urandom to create a file of random bytes, encrypted it with both rsyncrypto and GnuPG, and rsynced both of these to a remote system using rsync. I then modified the plaintext file, encrypted the file again, and synced the encrypted file with the remote system. In this case, I modified 6KB of data at an offset of 17KB into the file using dd and left all the other data intact. The final rsync commands show that the rsyncrypto-encrypted tree only needed to send 58,102 bytes, whereas the GnuPG-encrypted file required the entire file to be sent to the remote system:
$ cd ~/foodir $ rm -f * $ dd if=/dev/urandom of=testfile.random bs=1024 count=500 512000 bytes (512 kB) copied, 0.088045 s, 5.8 MB/s $ cd ~ $ rsyncrypto -r foodir foodir.rcrypto foodir.keys ~/rsyncrypto-keys/rckey.crt $ ls -l foodir.rcrypto/ -rw-r--r-- ... 502K 2008-01-08 19:59 testfile.random $ mkdir foodir.gpg $ gpg --gen-key ... $ mkdir foodir.gpg $ gpg --output foodir.gpg/testfile.random.gpg -e foodir/testfile.random $ ls -l foodir.gpg -rw-r--r-- 1 ... 501K 2008-01-08 20:07 testfile.random.gpg $ rsync -av foodir.rcrypto ben@v8tsrv:~ sent 513356 bytes received 48 bytes 342269.33 bytes/sec $ rsync -av foodir.gpg ben@v8tsrv:~ sent 513026 bytes received 48 bytes 342049.33 bytes/sec # # modify the input file starting at 17KB into the file for 6KB # $ dd if=/dev/urandom of=~/foodir/testfile.random bs=1024 count=6 seek=17 conv=notrunc $ ls -l testfile.random -rw-r--r-- 1 ... 500K 2008-01-08 20:17 testfile.random $ rsyncrypto -r foodir foodir.rcrypto foodir.keys ~/rsyncrypto-keys/rckey.crt $ gpg --output foodir.gpg/testfile.random.gpg -e foodir/testfile.random # # See how much gets sent # $ rsync -av foodir.rcrypto ben@v8tsrv:~ sent 58102 bytes received 4368 bytes 124940.00 bytes/sec $ rsync -av foodir.gpg ben@v8tsrv:~ sent 513024 bytes received 4368 bytes 1034784.00 bytes/sec
Using rsyncrypto with rsync, you can protect the files that you send to a remote system while allowing modified files to be sent in a bandwidth-efficient manner. There is a slight loss of security using rsyncrypto, because changes to the unencrypted file do not propagate throughout the entire encrypted file. When this security trade-off is acceptable, you can get much quicker bandwidth-friendly network syncs and still achieve good encryption on the files stored on the remote server.
If you wish to hide file names as well as their content on the remote server, you can use the --name-encrypt=map option to rsyncrypto, which stores a mapping from the original file name to a garbled random file name in a mapping file, and outputs files using only their garbled random file name in the encrypted directory tree.
Ben Martin has been working on filesystems for more than 10 years. He completed his Ph.D. and now offers consulting services focused on libferris, filesystems, and search solutions.
Is server or client processing better for charts and graphs?
By Colin Beckingham on February 01, 2008 (4:00:00 PM)
Webmasters are frequently required to serve up charts and graphs to clients. Part of the planning for such images involves a decision about whether to process the chart on the server or at the client end. Of course, it depends on the circumstances. There are costs and benefits to both approaches.
The generation of a chart at the server involves the creation of an image such as a .png file and then displaying this file as part of the delivered page. Prior to the image creation a script must set up the data points and the axis labels, switch on colours, create a legend, size the picture, and send it out.
Client-side processing sends the parameters and data points for processing by the client, often with JavaScript. However, actual delivery of the page involves sending not only the data and parameters, but also the required JavaScript and cascading stylesheet libraries to draw the final chart. The following table compares the two approaches.
Summary Comparison Server side Client side
Advantages
* Data remains confidential
* Chart can be copied and pasted as a separate unit
* If multiple copies of the chart are required, the image can be easily repeated
* Browser may be able to resize the image
* Images load progressively, visibly in the browser
* Data is available locally for reprocessing
* Once the .js and .css libraries are downloaded they can be re-used with no further download overhead
Disadvantages
* Identical data cannot be reprocessed without refreshing the page
* Heavier load on the server processor
* Server side image storage required
* Multiple identical images need to be separately drawn
* .js and .css libraries need to be part of the download package
* JavaScript must be available and activated on the client
* Taking a copy of a chart requires a screenshot
* Libraries load without screen activity, giving the appearance of slow loading
The open source world offers class libraries to provide either or both solutions. An example of a simple server-side library is libchart, and a client side library is Webfx Charting, both of which can be incorporated into a PHP script. Both projects provide samples of library output and example code. Each method probably requires approximately the same amount of coding effort. This article is not intended to compare the output of the two libraries, just to examine the relative advantages and disadvantages of the two approaches. Additionally, there are other methods such as Java which are not considered here.
The choice for webmasters comes down to two major components: the actual processing time, and the bandwidth required to deliver the final product. Generally the bandwidth component is easier to quantify. Some server managers place official quotas on bandwidth, but leave unofficial honour system quotas on processing time, reserving the right to reduce service to webmasters who place unfair burdens on processing time on shared servers. It is in the interests of webmasters to keep the server managers happy on both counts.
Download size
Focusing on the code necessary to draw the chart on the client machine, in the server side situation the code can be very short:

In the client side situation the minimum might be:
That's more than 800 characters, not counting spaces or data, about 450 of which is repeated for each chart to be displayed. Compared with 22 characters for the server side, that's not so trivial.
When it comes to data, the client-side is expecting an array of numbers. In the case that the array is short and consists of simple integers, the array is short and almost trivial. Now consider that the data might be an array of double-precision numbers such as [123456.7890,123456.7889,...,2.6]. A webmaster might be able to round down to integers before transmission, but this would lose some precision, which could be a problem when the data is to be re-used at the client end. Two hundred data points at 11 characters per number each plus the separation character gives an array character length of about 2,400.
In the case of one single chart we can expect something like:
Server side (bytes) Client side (bytes)
Basic html 22 Basic html 400
Libraries 0 Libraries (loaded once per page, if not cached) 66,000
Image 25,000 Image code 450
Data 0 Data (est., per image) 2,400
Total 25,022 Total 69,250
So the download for the client side will probably be larger for one image.
Now consider the case of five different charts on the same page:
Server side (bytes) Client side (bytes)
Basic html 100 Basic html 400
Libraries 0 Libraries (loaded once per page, if not cached) 66,000
Image 125,000 Image code 2,250
Data 0 Data (est., 5 images) 12,000
Total 125,100 Total 80,650
In the case of five charts, the client side download package is smaller. Clearly the issue of the size of the download depends entirely on the number of charts. While these numbers may appear insignificant, when a page is served up many thousands of times, the total becomes a threat to your bandwidth limit.
CPU load
Next we need to consider the issue of CPU load. The server-side class library uses the freetype and gd libraries. Some server managers consider these libraries to be processor-intensive, a critical issue in a shared server environment. You would expect that server-side processing would create a greater CPU load, but is that really so?
Using the top utility on a standalone single-processor server (2.8GHz Intel with hyperthreading and 2GB of RAM) I had the following experience. I set up a PHP script to generate a chart from 200 random data points using MySQL. One script output the charts as three separate 28KB .png images and the other drew three charts with JavaScript. Some of the CPU load can be attributed to MySQL, but since the database processing for each is the same, this should cancel out.
Results GD -> .png JavaScript
Page load seconds (*dialup) 15 3
CPU load Average increase of 10% for ~5 seconds Average increase of 3% for ~3 seconds
Download size (Source code + Other) 28K * 3 = 84K 8K + 66K = 74K
Conclusion
In summary, charting and graphing are important and informative assets. After all, a picture can be worth a thousand words. However, in a Web context, we need to balance whether it is better to send charts and graphs as pictures or the thousand words. Providing the data and then using JavaScript locally to draw the results not only gives a faster page load to the end user, but reduces the load on the server CPU, in this case by a factor of about three.
Colin Beckingham is a freelance programmer and writer in Eastern Ontario. He is currently working on a project that logs and charts the operation of a biomass burner using open source resources.
Webmasters are frequently required to serve up charts and graphs to clients. Part of the planning for such images involves a decision about whether to process the chart on the server or at the client end. Of course, it depends on the circumstances. There are costs and benefits to both approaches.
The generation of a chart at the server involves the creation of an image such as a .png file and then displaying this file as part of the delivered page. Prior to the image creation a script must set up the data points and the axis labels, switch on colours, create a legend, size the picture, and send it out.
Client-side processing sends the parameters and data points for processing by the client, often with JavaScript. However, actual delivery of the page involves sending not only the data and parameters, but also the required JavaScript and cascading stylesheet libraries to draw the final chart. The following table compares the two approaches.
Summary Comparison Server side Client side
Advantages
* Data remains confidential
* Chart can be copied and pasted as a separate unit
* If multiple copies of the chart are required, the image can be easily repeated
* Browser may be able to resize the image
* Images load progressively, visibly in the browser
* Data is available locally for reprocessing
* Once the .js and .css libraries are downloaded they can be re-used with no further download overhead
Disadvantages
* Identical data cannot be reprocessed without refreshing the page
* Heavier load on the server processor
* Server side image storage required
* Multiple identical images need to be separately drawn
* .js and .css libraries need to be part of the download package
* JavaScript must be available and activated on the client
* Taking a copy of a chart requires a screenshot
* Libraries load without screen activity, giving the appearance of slow loading
The open source world offers class libraries to provide either or both solutions. An example of a simple server-side library is libchart, and a client side library is Webfx Charting, both of which can be incorporated into a PHP script. Both projects provide samples of library output and example code. Each method probably requires approximately the same amount of coding effort. This article is not intended to compare the output of the two libraries, just to examine the relative advantages and disadvantages of the two approaches. Additionally, there are other methods such as Java which are not considered here.
The choice for webmasters comes down to two major components: the actual processing time, and the bandwidth required to deliver the final product. Generally the bandwidth component is easier to quantify. Some server managers place official quotas on bandwidth, but leave unofficial honour system quotas on processing time, reserving the right to reduce service to webmasters who place unfair burdens on processing time on shared servers. It is in the interests of webmasters to keep the server managers happy on both counts.
Download size
Focusing on the code necessary to draw the chart on the client machine, in the server side situation the code can be very short:

In the client side situation the minimum might be:
That's more than 800 characters, not counting spaces or data, about 450 of which is repeated for each chart to be displayed. Compared with 22 characters for the server side, that's not so trivial.
When it comes to data, the client-side is expecting an array of numbers. In the case that the array is short and consists of simple integers, the array is short and almost trivial. Now consider that the data might be an array of double-precision numbers such as [123456.7890,123456.7889,...,2.6]. A webmaster might be able to round down to integers before transmission, but this would lose some precision, which could be a problem when the data is to be re-used at the client end. Two hundred data points at 11 characters per number each plus the separation character gives an array character length of about 2,400.
In the case of one single chart we can expect something like:
Server side (bytes) Client side (bytes)
Basic html 22 Basic html 400
Libraries 0 Libraries (loaded once per page, if not cached) 66,000
Image 25,000 Image code 450
Data 0 Data (est., per image) 2,400
Total 25,022 Total 69,250
So the download for the client side will probably be larger for one image.
Now consider the case of five different charts on the same page:
Server side (bytes) Client side (bytes)
Basic html 100 Basic html 400
Libraries 0 Libraries (loaded once per page, if not cached) 66,000
Image 125,000 Image code 2,250
Data 0 Data (est., 5 images) 12,000
Total 125,100 Total 80,650
In the case of five charts, the client side download package is smaller. Clearly the issue of the size of the download depends entirely on the number of charts. While these numbers may appear insignificant, when a page is served up many thousands of times, the total becomes a threat to your bandwidth limit.
CPU load
Next we need to consider the issue of CPU load. The server-side class library uses the freetype and gd libraries. Some server managers consider these libraries to be processor-intensive, a critical issue in a shared server environment. You would expect that server-side processing would create a greater CPU load, but is that really so?
Using the top utility on a standalone single-processor server (2.8GHz Intel with hyperthreading and 2GB of RAM) I had the following experience. I set up a PHP script to generate a chart from 200 random data points using MySQL. One script output the charts as three separate 28KB .png images and the other drew three charts with JavaScript. Some of the CPU load can be attributed to MySQL, but since the database processing for each is the same, this should cancel out.
Results GD -> .png JavaScript
Page load seconds (*dialup) 15 3
CPU load Average increase of 10% for ~5 seconds Average increase of 3% for ~3 seconds
Download size (Source code + Other) 28K * 3 = 84K 8K + 66K = 74K
Conclusion
In summary, charting and graphing are important and informative assets. After all, a picture can be worth a thousand words. However, in a Web context, we need to balance whether it is better to send charts and graphs as pictures or the thousand words. Providing the data and then using JavaScript locally to draw the results not only gives a faster page load to the end user, but reduces the load on the server CPU, in this case by a factor of about three.
Colin Beckingham is a freelance programmer and writer in Eastern Ontario. He is currently working on a project that logs and charts the operation of a biomass burner using open source resources.
Subscribe to:
Posts (Atom)