TweetFollow Us on Twitter

June 96 - MPW Tips And Tricks: Scripted Text Editing

Mpw Tips And Tricks: Scripted Text Editing

Tim Maroney

The MPW Shell contains a full-strength, high-speed text editor with scripting capabilities. It's nothing to write love letters with, because it's targeted at the ASCII format of compiler source files, but it provides the power to automate complex and repetitive tasks in ASCII text. The key to the system lies in a few editing-related commands, together with its regular expressions and selection expressions.

REGULAR EXPRESSIONS

In the MPW Shell, any search command can take one of two kinds of arguments. The first is a plain string, which matches exactly its contents and nothing else, using a simple character-by-character match. The other is a regular expression, which is a pattern that can be recognized by a finite state machine. You can't parse programming languages with regular expressions, but you can use them to recognize many patterns, including wildcards, repeating sequences, and sets of characters. Regular expressions are bracketed with either slashes or backslashes, for searching forward or backward respectively. So, for instance, the regular expression \wombat\ would search backward from the current location for the string "wombat".

There are about 20 special constructs within regular expressions, all of which are cryptically described when you execute the command line "Help Patterns" within the MPW Shell. I'll mention some of the more useful ones here. The wildcard characters are the question mark (?) and the equivalence symbol (~, Option-X). The question mark matches any one character except the end of a line, while the equivalence symbol matches any number of such characters. For instance, /w?mb~t/ would match "wombat" as well as "wambiklort" and "wymbt", but not "wafkambiliot", nor "wkmb" at the end of a line. Restricted sets of symbols can be given in brackets; for instance, you can search for alphanumeric characters with the pattern [a-zA-Z0-9]. The reverse of a set can be specified with the "not" symbol (~, Option-L); for instance, /[~a-z]/ finds any character except a lowercase letter. The start of a line can be specified with the bullet symbol (*, Option-8) and the end of a line with the infinity symbol ([[infinity]], Option-5).

    These keyboard shortcuts are for American QWERTY keyboards. Other keyboards have different layouts. For instance, on a direct neural interface keyboard, think "blue wildebeest" and raise your right ear to type the bullet symbol.*
Repeating patterns can be specified in three ways. Following any pattern with a plus sign (+) means one or more instances of that pattern; for instance, the regular expression /[0-9]+/ would match any sequence of digits. An optional repeating pattern can be similarly specified with an asterisk (*), which means zero or more repetitions. The rarely seen double angle brackets can be used to specify exactly how many repetitions of a pattern are allowed. They're typed as Option-backslash (<<) and Option-Shift-backslash (>>) and enclose a single number to mean exactly that many repetitions, or two numbers separated by a comma to specify a minimum and maximum number of repetitions, or a single number followed by a comma to mean at least that many repetitions. For instance, the pattern /[a-zA-Z]<<3,7>>/ would find all strings composed of alphabetical characters and from three to seven letters long.

There are a number of ways of "escaping" special characters when you want to look for something that has special meaning within regular expressions, such as a question mark or plus sign. You can escape any character with the lowercase delta ([[partialdiff]], Option-D), or use single or double quotes to escape strings. To find the string "wombat+", for instance, you'd need to escape the plus sign: /wombat[[partialdiff]]+/.

Finally, one of the most useful constructs consists of a tagged regular expression. This allows you to associate a number between 0 and 9 with a pattern that's matched, referring to it later with the "registered" symbol (reg., Option-R) followed by a digit. This is very handy when you're doing replacements. For instance, you can replace any angle-bracketed string with a parenthesized string with the following command, which would turn "<wombat>" into "(wombat)":

Replace /<([~<>]*)reg.1>/ (reg.1)
This searches for any number of characters (except angle brackets) that are between angle brackets, assigns them the number 1, and then replaces the angle brackets with parentheses. Note that the syntax of tagged patterns requires the pattern to be parenthesized.

SELECTION EXPRESSIONS

Many editing commands (such as Replace) can take selection expressions as well as regular expressions. Selection expressions provide more ways to select text than the string matching provided by regular expressions. Common selection expressions include the following:
  • The bullet symbol, meaning the start of a file.

  • The infinity symbol, meaning the end of a file.

  • The current selection, denoted by [[section]] (Option-6). This might have been selected with the mouse or by a Find command. [[section]] by itself indicates the selection in the target window (which I'll explain later), while pathname:[[section]] means the selection in the file indicated by the pathname.

  • A line number, specified simply as a number.

  • The name of a marker, specified by the Mark command.

  • A range between two selection expressions, separated by a colon (:).
The above expressions require no special delimiters (they're not directional like regular expressions). Regular expressions are actually a kind of selection expression and are delimited by slash or backslash characters as usual.

Some character-skipping variants of these options are also provided, such as the position that's one character after the selection, denoted by following a selection expression with an uppercase delta ([[Delta]], Option-J). These are useful in dealing with context; for instance, you may want to select a string when it's followed by another character, but not include the following character in the selection. (An example is given later in the Subword script.) Text emitted by a program like a table generator may be in a known format, such as a columnar arrangement, in which case skipping a certain number of characters will take you to the selection you need.

Again, the MPW Shell will give you a terse summary of selection expressions when you execute the command line "Help Selections". I'm not going to list all the minor variants here, but feel free to while away the hours in rapturous contemplation of their mysteries on your own.

EDITING COMMANDS

The most common editing commands are two that you probably use already: Find and Replace. Dialogs that stand in for these commands are built into the MPW Shell and accessible from the Find menu. You can give any selection expression as a search pattern in either of these dialogs by clicking the Selection Expression radio button instead of the default Literal button. The same commands are the basis of most editing scripts. As tools, Find and Replace take a selection expression as their primary argument. Don't confuse Find and Search! The Search command puts out its results as text, while Find actually changes the selection. In addition, Search takes a pattern -- that is, a regular expression -- while Find takes any selection expression. For example, to go to the start of a file in a script, you could give the command "Find *", but not "Search *".

Find is the basic navigation command in most editing scripts. For instance, you can simulate the Select All command in the Edit menu like so:

Find *:[[infinity]]  # select from start to end of target
The commands File and Open, along with the variables Target and Active, determine the files your scripts will work on. "File" is actually an alias for the real command name, Target. The File command opens a file and makes it the target window -- the window behind the frontmost window. The target window is an important notion in MPW. It exists so that you can use the Worksheet window to type commands that affect another window; since the Worksheet would be in front, the window being affected would need to be behind the Worksheet. During scripting, you may prefer to use the Open command, which opens a file and makes it the frontmost window. The target window is referred to as {Target} in scripts, while the frontmost window is called {Active}. Editing commands work on the target window if you don't specify a window explicitly.

The Line command may also be used for navigation: it selects the numbered line in the target window and then brings that window to the front. You probably know this command already if you use compilers in the MPW Shell, since they put out error messages in this form:

File "gwork.c"; Line 418 # Syntax error
Executing this command takes you to the line in your code where the error was detected.

The Position command returns the current position in the target window, as a line number, a character range, or both. The position could be saved to a variable for later use as follows, using the backquote mechanism to execute a command and insert its output inline:

Set SavedLineNumber `Position -l`
There are dozens of commands pertaining to text editing in the MPW scripting language. Help on all of them is available in the MPW Shell. The usual Macintosh text-editing menu commands are available in the MPW scripting language, including New, Open, Close, Save, Revert, Print, and the standard Edit menu commands.

StreamEdit is a standalone editing tool that's rich and strange enough to deserve its own co-->umn. It's a structured search and replacement language based on the UNIXreg. command sed.

Some simpler standalone editing tools are provided. Sort has a rich function set and can be used for many text-editing tasks. Canon takes a file of search and replace strings and applies them to a file. It's used to automate terminology changes, such as the work that was done to make the Mac OS API use fewer acronyms and abbreviations when the new Inside Macintosh books were written. Translate, like the UNIX command tr, maps characters onto other characters.

Text indentation can be handled with four tools: Adjust, Align, Entab, and Format. Adjust shifts a line to the right or left by a specified number of spaces. Align sets the margin of a range of selected lines to the margin of the first selected line. Entab converts runs of spaces to tabs, and Format sets the column width used for tabs in a text document, as well as other settings like font and size. (These settings are saved in a resource in the file, which many ASCII text editors can recognize.)

Text-editing scripts often create temporary files, split single files into multiple files, and perform other file-related tasks. MPW provides commands to help you manage files. It has commands corresponding to almost all Finder operations, such as Duplicate, Move, Delete, and NewFolder. There are also some specialized file commands: FileDiv splits a file into multiple files based on a byte or line count or on embedded form feed characters inserted during a previous editing pass; Catenate does the opposite, joining files together.

A text-editing script often takes search and substitution text as parameters on the command line. A few commands related to parameters are worth a quick mention here. Echo is handy for concatenating parameters with other text. Quote is similar to Echo but adds quote marks as needed to preserve the word breaks in its parameters. MPW scripting requires quotes around any string that is meant to be a single parameter but contains spaces (which would break the string into multiple parameters). Echo puts out its arguments in a way that allows them to be broken up, while Quote preserves the original word breaks by inserting quotes.

Echo "Richard Loves Pat"
Richard Loves Pat
Quote "Bill Loves Everyone"
'Bill Loves Everyone'

AN EXAMPLE SCRIPT

Here's a script I've found useful for some years. It's called Subword and it replaces a word by another string everywhere it occurs in the target window.
Set Sep "[~a-zA-Z_0-9]"  # word separators
Find * "{Target}"  # start at top of file
Replace -c [[infinity]] [[partialdiff]]
   "[[Delta]]/{Sep}{1}{Sep}/!1:[[Delta]]/{Sep}/" [[partialdiff]]
   "{2}" "{Target}"
The selection in this Replace command is probably about as clear as the U.S. tax code, so allow me to explain. The [[Delta]] means one character before the selection. The !1 means one character past the selection. The colon denotes everything between the selections (inclusively). So this pattern says, in a nutshell, select the pattern in the first parameter ({1}) when it's bracketed by separators, but exclude the separators.

Normally I don't use this script directly. I incorporate it into other scripts as a utility. The bulk of the work of converting between similar languages like Pascal and C can be done by an editing script, for example. Subword can be used to convert keywords, as could Canon. I use another script which is essentially Subword without the separators for changing symbols like equality operators.

Scripts to preconvert between Pascal and C can be found on this issue's CD. They don't generate compiler-ready text, but I've found that they facilitate a manual conversion at the rate of hundreds of lines per hour, allowing source bases in the thousands of lines to be accurately translated in a day or three. So the next time you're faced with a dull text-processing task, look over the tools MPW gives you, and see whether you can save yourself a few days of tedious manual labor!

TIM MARONEY recently changed his Apple badge color from green to white: he's gone from contract programming to a technical leadership role developing user interface software. Tim entertains himself in a variety of ways, such as straining his surgically altered eyeballs on the small print of obscure footnotes and collectible trading card games, and contorting his limbs in yogic asanas. He designed the iron crystal that now resides at the core of the earth and contributed significant ideas to the original (now obsolete) implementation of Planck-scale gravitational phenomena in the universe.*

Thanks to Dave Evans, Scott Fraser, Arno Gourdol, and Alex McKale for reviewing this column.*

 

Community Search:
MacTech Search:

Software Updates via MacUpdate

Coda 2.5.11 - One-window Web development...
Coda is a powerful Web editor that puts everything in one place. An editor. Terminal. CSS. Files. With Coda 2, we went beyond expectations. With loads of new, much-requested features, a few surprises... Read more
Bookends 12.5.7 - Reference management a...
Bookends is a full-featured bibliography/reference and information-management system for students and professionals. Access the power of Bookends directly from Mellel, Nisus Writer Pro, or MS Word (... Read more
Maya 2016 - Professional 3D modeling and...
Maya is an award-winning software and powerful, integrated 3D modeling, animation, visual effects, and rendering solution. Because Maya is based on an open architecture, all your work can be scripted... Read more
RapidWeaver 6.2.3 - Create template-base...
RapidWeaver is a next-generation Web design application to help you easily create professional-looking Web sites in minutes. No knowledge of complex code is required, RapidWeaver will take care of... Read more
MacFamilyTree 7.5.2 - Create and explore...
MacFamilyTree gives genealogy a facelift: it's modern, interactive, incredibly fast, and easy to use. We're convinced that generations of chroniclers would have loved to trade in their genealogy... Read more
Paragraphs 1.0.1 - Writing tool just for...
Paragraphs is an app just for writers. It was built for one thing and one thing only: writing. It gives you everything you need to create brilliant prose and does away with the rest. Everything in... Read more
BlueStacks App Player 0.9.21 - Run Andro...
BlueStacks App Player lets you run your Android apps fast and fullscreen on your Mac. Version 0.9.21: Note: Now requires OS X 10.8 or later running on a 64-bit Intel processor. Initial stable... Read more
Tweetbot 2.0.2 - Popular Twitter client....
Tweetbot is a full-featured OS X Twitter client with a lot of personality. Whether it's the meticulously-crafted interface, sounds and animation, or features like multiple timelines and column views... Read more
Apple iBooks Author 2.3 - Create and pub...
Apple iBooks Author helps you create and publish amazing Multi-Touch books for iPad. Now anyone can create stunning iBooks textbooks, cookbooks, history books, picture books, and more for iPad. All... Read more
NeoOffice 2014.12 - Mac-tailored, OpenOf...
NeoOffice is a complete office suite for OS X. With NeoOffice, users can view, edit, and save OpenOffice documents, PDF files, and most Microsoft Word, Excel, and PowerPoint documents. NeoOffice 3.x... Read more

Rage of Bahamut is Giving Almost All of...
The App Store isn't what it used to be back in 2012, so it's not unexpected to see some games changing their structures with the times. Now we can add Rage of Bahamut to that list with the recent announcement that the game is severely cutting back... | Read more »
Adventures of Pip (Games)
Adventures of Pip 1.0 Device: iOS iPhone Category: Games Price: $4.99, Version: 1.0 (iTunes) Description: ** ONE WEEK ONLY — 66% OFF! *** “Adventures of Pip is a delightful little platformer full of charm, challenge and impeccable... | Read more »
Divide By Sheep - Tips, Tricks, and Stre...
Who would have thought splitting up sheep could be so involved? Anyone who’s played Divide by Sheep, that’s who! While we’re not about to give you complete solutions to everything (because that’s just cheating), we will happily give you some... | Read more »
NaturalMotion and Zynga Have Started Tea...
An official sequel to 2012's CSR Racing is officially on the way, with Zynga and NaturalMotion releasing a short teaser trailer to get everyone excited. Well, as excited as one can get from a trailer with no gameplay footage, anyway. [Read more] | Read more »
Grab a Friend and Pick up Overkill 3, Be...
Overkill 3 is a pretty enjoyable third-person shooter that was sort of begging for some online multiplayer. Fortunately the begging can stop, because its newest update has added an online co-op mode. [Read more] | Read more »
Scanner Pro's Newest Update Adds Au...
Scanner Pro is one of the most popular document scanning apps on iOS, thanks in no small part to its near-constant updates, I'm sure. Now we're up to update number six, and it adds some pretty handy new features. [Read more] | Read more »
Heroki (Games)
Heroki 1.0 Device: iOS Universal Category: Games Price: $7.99, Version: 1.0 (iTunes) Description: CLEAR THE SKIES FOR A NEW HERO!The peaceful sky village of Levantia is in danger! The dastardly Dr. N. Forchin and his accomplice,... | Read more »
Wars of the Roses (Games)
Wars of the Roses 1.0 Device: iOS Universal Category: Games Price: $4.99, Version: 1.0 (iTunes) Description: | Read more »
TapMon Battle (Games)
TapMon Battle 1.0 Device: iOS Universal Category: Games Price: $.99, Version: 1.0 (iTunes) Description: It's time to battle!Tap! Tap! Tap! Try tap a egg to hatch a Tapmon!Do a battle with another tapmons using your hatched tapmons! *... | Read more »
Alchemic Dungeons (Games)
Alchemic Dungeons 1.0 Device: iOS Universal Category: Games Price: $.99, Version: 1.0 (iTunes) Description: ### Release Event! ### 2.99$->0.99$ for limited time! ### Roguelike Role Playing Game! ### Alchemic Dungeons is roguelike... | Read more »

Price Scanner via MacPrices.net

15-inch 2.5GHz Retina MacBook Pro on sale for...
Amazon.com has the 15″ 2.5GHz Retina MacBook Pro on sale for $2274 including free shipping. Their price is $225 off MSRP, and it’s the lowest price available for this model. Read more
Logo Pop Free Vector Logo Design App For OS X...
128bit Technologies has released of Logo Pop Free 1.2 for Mac OS X, a vector based, full-fledged, logo design app available exclusively on the Mac App Store for the agreeable price of absolutely free... Read more
21-inch 1.4GHz iMac on sale for $999, save $1...
B&H Photo has new 21″ 1.4GHz iMac on sale for $999 including free shipping plus NY sales tax only. Their price is $100 off MSRP. Best Buy has the 21″ 1.4GHz iMac on sale for $999.99 on their... Read more
16GB iPad mini 3 on sale for $339, save $60
B&H Photo has the 16GB iPad mini 3 WiFi on sale for $339 including free shipping plus NY tax only. Their price is $60 off MSRP. Read more
Save up to $40 on iPad Air 2, NY tax only, fr...
B&H Photo has iPad Air 2s on sale for up to $40 off MSRP including free shipping plus NY sales tax only: - 16GB iPad Air 2 WiFi: $489 $10 off - 64GB iPad Air 2 WiFi: $559 $40 off - 128GB iPad Air... Read more
Apple Releases OS X 10.10.4 With WIFi Fix, iO...
On Tuesday, Apple released final versions of OS X 10.10.4 and iOS 8.4, as well as updates for the Safari browser for OS X Yosemite, Mavericks, and Mountain Lion. The OS X 10.10.4 update focuses on... Read more
Dual-Band High-Gain Antennas for Home Wi-Fi N...
Linksys has announced what it claims are the first dual-band, omni-directional high-gain antennas for the consumer market. The new Linksys high-gain antennas available in a 2- and 4-pack (WRT004ANT... Read more
Apple refurbished 2014 15-inch Retina MacBook...
The Apple Store has Apple Certified Refurbished 2014 15″ 2.2GHz Retina MacBook Pros available for $1609, $390 off original MSRP. Apple’s one-year warranty is included, and shipping is free. They have... Read more
Clearance 2014 MacBook Airs available for up...
Adorama has 2014 MacBook Airs on sale for up to $301 off original MSRP including NY + NJ sales tax and free shipping: - 11″ 256GB MacBook Air: $798 $301 off original MSRP - 13″ 128GB MacBook Air: $... Read more
5K iMacs on sale for $100 off MSRP, free ship...
B&H Photo has the new 27″ 3.3GHz 5K iMac on sale for $1899.99 including free shipping plus NY tax only. Their price is $100 off MSRP. They have the 27″ 3.5GHz 5K iMac on sale for $2199, also $100... Read more

Jobs Board

*Apple* Solutions Consultant - Retail Sales...
**Job Summary** As an Apple Solutions Consultant (ASC) you are the link between our customers and our products. Your role is to drive the Apple business in a retail Read more
*Apple* Fulfillment Operations Execution Ana...
**Job Summary** The AMR Apple Fulfillment Operations Team is seeking a talented team player to drive the Apple Online Store (AOS) fulfillment performance to ensure a Read more
Localization Producer - *Apple* HR and Reta...
…project manager to support the Retail Globalization team. You will participate in Apple exponential inte ational growth and drive global project initiatives for the Read more
*Apple* Online Store UAT Lead - Apple (Unite...
**Job Summary** The Apple Online Store is a fast paced and ever evolving business environment. A UAT lead in this organization is able to have a direct impact on one of Read more
Senior Payments Security Manager - *Apple*...
**Job Summary** Apple , Inc. is looking for a highly motivated, innovative and hands-on senior payments security manager to join the Apple Pay security team. You will Read more
All contents are Copyright 1984-2011 by Xplain Corporation. All rights reserved. Theme designed by Icreon.