Tabla de Contenidos

AWK

Basics I

$1 Reference first column
$2 Reference second column
; Char to separate two actions
print Print current record line
$0 Reference current record line
^ Match beginning of field
\~ Match opterator
\!\~ Do not match operator
\-F Command line option to specify input field delimiter
BEGIN Denotes block executed once at start
END Denotes block executed once at end
str1 str2 Concat str1 and str2

Variables

FS Field separator of input file (default whitespace)
NF Number of fields in current record
NR Line number of the current record
FILENAME Reference current input file
FNR Reference number of the current record relative to current input file
OFS Field separator of the outputted data (default whitespace)
ORS Record separator of the outputted data (default newline)
RS Record separator of input file (default newline)
CONVFMT Conversion format used when converting numbers (default %.6g)
SUBSEP Separates multiple subscripts (default 034)
OFMT Output format for numbers (default %.6g)
ARGC Argument count, assignable
ARGV Argument array, assignable
ENVIRON Array of environment variables

Functions

index(s,t) Position in string s where string t occurs, 0 if not found
length(s) Length of string s (or $0 if no arg)
rand() Random number between 0 and 1
substr(s,i,l) Return l len-char substring of s that begins at index i (counted from 1)
srand() Set seed for rand and return previous seed
int(x) Truncate x to integer value
split(s,a,fs) Split string s into array a split by fs, returning length of a
match(s,r) Position in string s where regex r occurs, or 0 if not found
sub(r,t,s) Substitute t for first occurrence of regex r in string s (or $0 if s not given)
gsub(r,t,s) Substitute t for all occurrences of regex r in string s
system(cmd) Execute cmd and return exit status
tolower(s) String s to lowercase
toupper(s) String s to uppercase
getline() Set $0 to next input record from current input file.

OneLiners

Execute action for matched pattern ‘pattern’ on file ‘file’

  awk /pattern/ {action} file

Print first field for each record in file

  awk '{print $1}' file

Print only lines that match regex in file

  awk '/regex/' file

Print only lines that do not match regex in file

  awk '!/regex/' file

Print any line where field 2 is equal to “foo” in file

  awk '$2 == "foo"' file

Print lines where field 2 is NOT equal to “foo” in file

  awk '$2 != "foo"' file

Print line if field 1 matches regex in file

  awk '$1 ~ /regex/' file

Print line if field 1 does NOT match regex in file

  awk '$1 !~ /regex/' file

Print first field for each record in file excluding the first record

  awk 'NR!=1{print $1}' file

Count lines in file

  awk 'END{print NR}' file

Print total number of lines that contain foo

  awk '/foo/{n++}; END {print n+0}' file

Print total number of fields in all lines

  awk '{total=total+NF};END{print total}' file

Print line immediately after regex, but not line containing regex in file

  awk '/regex/{getline;print}' file

Print lines with more than 32 characters in file

  awk 'length > 32' file

Print line number 12 of file

  awk 'NR==12' file

Precede each line by its line number FOR THAT FILE (left alignment). Using a tab ( instead of space will preserve margins.

  awk '{print FNR "\t" $0}' files*

Precede each line by its line number FOR ALL FILES TOGETHER, with tab.

  awk '{print NR "\t" $0}' files*

Print diagram of user/groups

  awk 'BEGIN{FS=":"; print "digraph{"}{split($4, a, ",");for (i in a) printf "\"%s\" [shape=box];\n\"%s\" -> \"%s\";\n", $1, a[i], $1}END{print "}"}' /etc/group|display

print the sums of the fields of every line

  awk '{s=0; for (i=1; i<=NF; i++) s=s+$i; print s}'

add all fields in all lines and print the sum

  awk '{for (i=1; i<=NF; i++) s=s+$i}; END{print s}'

print every line after replacing each field with its absolute value

  awk '{for (i=1; i<=NF; i++) if ($i < 0) $i = -$i; print }'
  awk '{for (i=1; i<=NF; i++) $i = ($i < 0) ? -$i : $i; print }'

print the total number of fields (“words”) in all lines

  awk '{ total = total + NF }; END {print total}' file

print the total number of lines that contain “Beth”

  awk '/Beth/{n++}; END {print n+0}' file

print the largest first field and the line that contains it Intended for finding the longest string in field \#1

  awk '$1 > max {max=$1; maxline=$0}; END{ print max, maxline}'

print the number of fields in each line, followed by the line

  awk '{ print NF ":" $0 } '

print the last field of each line

  awk '{ print $NF }'

print every line with more than 4 fields

  awk 'NF > 4'

print every line where the value of the last field is \> 4

  awk '$NF > 4'

create a string of a specific length (e.g., generate 513 spaces)

  awk 'BEGIN{while (a++<513) s=s " "; print s}'

double space a file

  awk '1;{print ""}'
  awk 'BEGIN{ORS="\n\n"};1'

double space a file which already has blank lines in it. Output file should contain no more than one blank line between lines of text. NOTE: On Unix systems, DOS lines which have only CRLF () are often treated as non-blank, and thus ‘NF’ alone will return TRUE.

  awk 'NF{print $0 "\n"}'

triple space a file

  awk '1;{print "\n"}'

IN UNIX ENVIRONMENT: convert DOS newlines (CR/LF) to Unix format

  awk '{sub(/\r$/,"")};1'   # assumes EACH line ends with Ctrl-M

IN UNIX ENVIRONMENT: convert Unix newlines (LF) to DOS format

  awk '{sub(/$/,"\r")};1'

IN DOS ENVIRONMENT: convert Unix newlines (LF) to DOS format

  awk 1

IN DOS ENVIRONMENT: convert DOS newlines (CR/LF) to Unix format Cannot be done with DOS versions of awk, other than gawk

  gawk -v BINMODE="w" '1' infile >outfile

delete leading whitespace (spaces, tabs) from front of each line aligns all text flush left

  awk '{sub(/^[ \t]+/, "")};1'

delete trailing whitespace (spaces, tabs) from end of each line

  awk '{sub(/[ \t]+$/, "")};1'

delete BOTH leading and trailing whitespace from each line

  awk '{gsub(/^[ \t]+|[ \t]+$/,"")};1'
  awk '{$1=$1};1'           # also removes extra space between fields

insert 5 blank spaces at beginning of each line (make page offset)

awk '{sub(/^/, "     ")};1'

align all text flush right on a 79-column width

  awk '{printf "%79s\n", $0}' file*

center all text on a 79-character width

  awk '{l=length();s=int((79-l)/2); printf "%"(s+l)"s\n",$0}' file*

substitute (find and replace) “foo” with “bar” on each line

  awk '{sub(/foo/,"bar")}; 1'           # replace only 1st instance
  gawk '{$0=gensub(/foo/,"bar",4)}; 1'  # replace only 4th instance
  awk '{gsub(/foo/,"bar")}; 1'          # replace ALL instances in a line

substitute “foo” with “bar” ONLY for lines which contain “baz”

  awk '/baz/{gsub(/foo/, "bar")}; 1'

substitute “foo” with “bar” EXCEPT for lines which contain “baz”

  awk '!/baz/{gsub(/foo/, "bar")}; 1'

change “scarlet” or “ruby” or “puce” to “red”

  awk '{gsub(/scarlet|ruby|puce/, "red")}; 1'

reverse order of lines (emulates “tac”)

  awk '{a[i++]=$0} END {for (j=i-1; j>=0;) print a[j--] }' file*

if a line ends with a backslash, append the next line to it (fails if there are multiple lines ending with backslash…)

  awk '/\\$/ {sub(/\\$/,""); getline t; print $0 t; next}; 1' file*

print and sort the login names of all users

  awk -F ":" '{print $1 | "sort" }' /etc/passwd

print the first 2 fields, in opposite order, of every line

  awk '{print $2, $1}' file

switch the first 2 fields of every line

  awk '{temp = $1; $1 = $2; $2 = temp}' file

print every line, deleting the second field of that line

  awk '{ $2 = ""; print }'

print in reverse order the fields of every line

  awk '{for (i=NF; i>0; i--) printf("%s ",$i);print ""}' file

concatenate every 5 lines of input, using a comma separator between fields

  awk 'ORS=NR%5?",":"\n"' file

count lines (emulates “wc -l”)

  awk 'END{print NR}'

print first 10 lines of file (emulates behavior of “head”)

  awk 'NR < 11'

print first line of file (emulates “head -1”)

  awk 'NR>1{exit};1'

print the last 2 lines of a file (emulates “tail -2”)

  awk '{y=x "\n" $0; x=$0};END{print y}'

print the last line of a file (emulates “tail -1”)

  awk 'END{print}'

print only lines which match regular expression (emulates “grep”)

  awk '/regex/'

print only lines which do NOT match regex (emulates “grep -v”)

  awk '!/regex/'

print any line where field \#5 is equal to “abc123”

  awk '$5 == "abc123"'

print only those lines where field \#5 is NOT equal to “abc123” This will also print lines which have less than 5 fields.

  awk '$5 != "abc123"'
  awk '!($5 == "abc123")'

matching a field against a regular expression

  awk '$7  ~ /^[a-f]/'    # print line if field #7 matches regex
  awk '$7 !~ /^[a-f]/'    # print line if field #7 does NOT match regex

print the line immediately before a regex, but not the line containing the regex

  awk '/regex/{print x};{x=$0}'
  awk '/regex/{print (NR==1 ? "match on line 1" : x)};{x=$0}'

print the line immediately after a regex, but not the line containing the regex

  awk '/regex/{getline;print}'

grep for AAA and BBB and CCC (in any order on the same line)

  awk '/AAA/ && /BBB/ && /CCC/'

grep for AAA and BBB and CCC (in that order)

  awk '/AAA.*BBB.*CCC/'

print only lines of 65 characters or longer

  awk 'length > 64'

print only lines of less than 65 characters

  awk 'length < 64'

print section of file from regular expression to end of file

  awk '/regex/,0'
  awk '/regex/,EOF'

print section of file based on line numbers (lines 8-12, inclusive)

  awk 'NR==8,NR==12'

print line number 52

  awk 'NR==52'
  awk 'NR==52 {print;exit}'          # more efficient on large files

print section of file between two regular expressions (inclusive)

  awk '/Iowa/,/Montana/'             # case sensitive

delete ALL blank lines from a file (same as “grep ‘.’”)

  awk NF
  awk '/./'

remove duplicate, consecutive lines (emulates “uniq”)

  awk 'a !~ $0; {a=$0}'

remove duplicate, nonconsecutive lines

  awk '!a[$0]++'                     # most concise script
  awk '!($0 in a){a[$0];print}'      # most efficient script

number each line of a file (number on left, right-aligned) Double the percent signs if typing from the DOS command prompt.

  awk '{printf("%5d : %s\n", NR,$0)}'

number each line of file, but only print numbers if line is not blank Remember caveats about Unix treatment of mentioned above)

  awk 'NF{$0=++a " :" $0};1'
  awk '{print (NF? ++a " :" :"") $0}'

Filtrar columnas

Sirve para filtrar por columnas en una cadena de texto:

IMPRIMIR ÚLTIMA COLUMNA

Si queremos imprimir la última columna con awk, utilizaríamos $NF:

Otro ejemplo de cómo imprimir la última columna con awk:

CONECTAR COLUMNAS También podemos conectar unas columnas con otras de la siguiente forma, por ejemplo agregando flechas:

ESTABLECER UN DELIMITADOR

También podemos establecer un delimitador para que nos cuente a partir de un determinado patrón, por ejemplo de un punto:

Otro ejemplo: