Biochemistry EIGHTH EDITION Jeremy M. Berg Tymoczko Gregory J. Gatto, Jr. John L. Lubert Stryer Publisher: Kate Ahr Parker Senior Acquisitions Editor: Lauren Schultz Developmental Editor: Irene Pech Editorial Assistants: Shannon Moloney and Nandini Ahuja Senior Project Editor: Denise Showers with Sherrill Redd Manuscript Editors: Irene Vartanoff and Mercy Heston Cover and Interior Design: Vicki Tomaselli Illustrations: Jeremy Berg with Network Graphics, Gregory J. Gatto, Jr. Illustration Coordinator: Janice Donnola Photo Editor: Christine Buese Photo Researcher: Jacquelyn Wong Production Coordinator: Paul Rohloff Executive Media Editor: Amanda Dunning Media Editor: Donna Brodman Executive Marketing Manager: Sandy Lindelof Composition: Aptara®, Inc. Printing and Binding: RR Donnelley Library of Congress Control Number: 2014950359 Gregory J. Gatto, Jr., is an employee of GlaxoSmithKline (GSK), which has not supported or funded this work in any way. Any views expressed herein do not necessarily represent the views of GSK. ISBN-13: 978-1-4641-2610-9 ISBN-10: 1-4641-2610-0 ©2015, 2012, 2007, 2002 by W. H. Freeman and Company; © 1995, 1988, 1981, 1975 by Lubert Stryer All rights reserved Printed in the United States of America First printing W. H. Freeman and Company 41 Madison Avenue New York, NY 10010 www.whfreeman.com To our teachers and our students Professor of Computational and Systems Biology and Pittsburgh Foundation Professor and Director of the Institute for Personalized Medicine. He JEREMY M. BERG received his B.S. and M.S. served as President of the American Society for degrees in Chemistry from Stanford (where he did Biochemistry and Molecular Biology from 2011– research with Keith Hodgson and Lubert Stryer) 2013. He is a Fellow of the American Association and his Ph.D. in Chemistry from Harvard with for the Advancement of Science and a member of Richard Holm. He then completed a postdoctoral the Institute of Medicine of the National Academy fellowship with Carl Pabo in Biophysics at Johns of Sciences. He received the American Chemical Hopkins University School of Medicine. He was an Society Award in Pure Chemistry (1994) and the Assistant Professor in the Department of Eli Lilly Award for Fundamental Research in Chemistry at Johns Hopkins from 1986 to 1990. Biological Chemistry (1995), was named Maryland He then moved to Johns Hopkins University Outstanding Young Scientist of the Year (1995), School of Medicine as Professor and Director of received the Harrison Howe Award (1997), and the Department of Biophysics and Biophysical received public service awards from the Chemistry, where he remained until 2003. He then Biophysical Society, the American Society for became Director of the National Institute of Biochemistry and Molecular Biology, the American General Medical Sciences at the National Chemical Society, and the American Society for Institutes of Health. In 2011, he moved to the Cell Biology. He also received numerous teaching University of Pittsburgh where he is now ABOUT THE AUTHORS awards, including the W. Barry Wood Teaching Award (selected by medical students), the Graduate Student Teaching Award, and the Professor’s Teaching Award for the Preclinical Sciences. He is coauthor, with Stephen J. Lippard, of the textbook Principles of Bioinorganic Chemistry. JOHN L. TYMOCZKO is Towsley Professor of Biology at Carleton College, where he has taught since 1976. He currently teaches Biochemistry, Biochemistry Laboratory, Oncogenes and the Molecular Biology of Cancer, and Exercise Biochemistry and coteaches an introductory course, Energy Flow in Biological Systems. Professor Tymoczko received his B.A. from the University of Chicago in 1970 and his Ph.D. in Biochemistry from the University of Chicago with Shutsung Liao at the Ben May Institute for Cancer Research. He then had a postdoctoral position with Hewson Swift of the Department of Biology at the University of Chicago. The focus of his research has been on steroid recep tors, ribonucleoprotein particles, and proteolytic processing enzymes. GREGORY J. GATTO, JR., received his A.B. degree in Chemistry from Princeton University, iv PREFACE where he worked with Martin F. Semmelhack and was awarded the Everett S. Wallis Prize in Organic Chemistry. In 2003, he received his M.D. and Ph.D. degrees from the Johns Hopkins University School of Medicine, where he studied the structural biology of peroxisomal targeting signal recognition with Jeremy M. Berg and received the Michael A. Shanoff Young Investigator Research Award. He completed a postdoctoral fellowship in 2006 with Christopher T. Walsh at Harvard Medical School, where he studied the biosynthesis of the macrolide immunosuppres sants. He is currently a Senior Scientific Investigator in the Heart Failure Discovery Performance Unit at GlaxoSmithKline. LUBERT STRYER is Winzer Professor of Cell Biology, Emeritus, in the School of Medicine and Professor of Neurobiology, Emeritus, at Stanford University, where he has been on the faculty since 1976. He received his M.D. from Harvard Medical School. Professor Stryer has received many awards for his research on the interplay of light and life, including the Eli Lilly Award for Fundamental Research in Biological Chemistry, the Distinguished Inventors Award of the Intellectual Property Owners’ Association, and election to the National Academy of Sciences and the American Philosophical Society. He was awarded the National Medal of Science in 2006. The publication of his first edition of Biochemistry in 1975 transformed the teaching of biochemistry. to help students see main points without the distraction of excess detail. • Physiological relevance. It has always been our goal to help students connect biochemistry to their own lives on a variety of scales. Pathways and pro cesses are presented in a physiological context so For several generations of students and teachers, Biochemistry has been an invaluable resource, pre senting the concepts and details of molecular structure, metabolism, and laboratory techniques in a streamlined and engaging way. Biochemistry’s success in helping students learn the subject for the first time is built on a number of hallmark features: • Clear writing and simple illustrations. The lan guage of biochemistry is made as accessible as possi ble for students learning the subject for the first time. To complement the straightforward language and organization of concepts in the text, figures illustrate a single concept at a time students can see how biochemistry works in the body and under different conditions, and Clinical Application sections in every chapter show students how the concepts they are studying impact human health. The eighth edition includes a number of new Clinical Application sections based on recent dis coveries in biochemistry and health. (For a full list, see p. xi) • Evolutionary perspective. Discussions of evolution are woven into the narrative of the text, just as evolu tion shapes every pathway and molecular structure described in the text. Molecular Evolution sections highlight important milestones in the evolution of life as a way to provide context for the processes and molecules being discussed. (For a full list, see p. x) • Problem-solving practice. Every chapter of Biochemistry provides numerous opportunities for students to practice problem-solving skills and apply the concepts described in the text. End-of-chapter interpretation structures. All 0% problems ask stu molecular structures 1.0 dents to draw in the book, with few conclusions from exceptions, have 0.9 50% data taken from real been selected and 0.8 research papers; rendered by Jeremy 0.7 AB and chapter Berg and Gregory integration problems Gatto to emphasize require students to the aspect of connect concepts structure most impor from across tant to the topic at chapters. Further hand. Students are 0% problem-solving introduced to problems are divided practice is pro vided realistic renderings into three categories online, on the of molecules through to address different Biochemistry a molecular model problem-solving LaunchPad. (For “primer” in the skills: Mechanism 100% 50% more details on appendices to prob lems ask LaunchPad Chapters 1 and 2 so students to suggest 100% resources, see p. they are wellor describe a viii) equipped to chemical recognize and • A variety of mechanism; Data interpret molecular Light aerobic effort the structures throughout the book. Figure legends Maximal aerobic effort direct students explicitly to the key features of a Figure 27.12 An idealized representation of fuels use as a function of model, and often include PDB numbers so the aerobic exercise intensity. (A) With increased exercise intensity, the reader can access the file used in generating use of fats as fuels falls as the utilization of glucose increases. (B) the structure from the Protein Data Bank The respiratory quotient (RQ) measures the alteration in fuel use. website (www.pdb.org). Students r b o h y d r a t e u t n i o l it i a z z a ili t t i u o n t a F QR C a v Time (sec) vi Preface (A) ) m n ( (B) n o 800 i t i Myosin V dimer s o P 1200 1000 600 400 Catalytic domain 74 nm 200 0 10 3020 40 50 60 70 80 90 100 110 Actin Figure 9.48 Single molecule motion. (A) A trace of the position of a single dimeric myosin V molecule as it moves across a surface coated with actin filaments. (B) A model of how the dimeric molecule moves in discrete steps with an average size of 74 6 5 nm. [Data from A. Yildiz et al., Science 300(5628)2061–2065, 2003.] can explore molecular structures further online through the Living Figures, in which they can rotate 3D models of molecules and view alternative renderings. In this revision of Biochemistry, we focused on build ing on the strengths of the previous editions to present biochemistry in an even more clear and streamlined manner, as well as incorporating exciting new advances from the field. Throughout the book, we have updated explanations of basic concepts and bolstered them with examples from new research. Some new topics that we present in the eighth edition include: • Environmental factors that influence human biochemistry (Chapter 1) • Genome editing (Chapter 5) • Horizontal gene transfer events that may explain unex pected branches of the evolutionary tree (Chapter 6) • Penicillin irreversibly inactivating a key enzyme in bacterial cell-wall synthesis (Chapter 8) • Scientists watching single molecules of myosin move (Chapter 9) • Glycosylation functions in nutrient sensing (Chapter 11) • The structure of a SNARE complex (Chapter 12) • The mechanism of ABC transporters (Chapter 13) • The structure of the gap junction (Chapter 13) • The structural basis for activation of the badrenergic receptor (Chapter 14) • Excessive fructose consumption can lead to patho logical conditions (Chapter 16) • Alterations in the glycolytic pathway by cancer cells (Chapter 16) • Regulation of mitochondrial ATP synthase (Chapter 18) • Control of chloroplast ATP synthase (Chapter 19) • Activation of rubisco by rubisco activase (Chapter 20) Figure 12.39 SNARE complexes initiate membrane fusion. The SNARE protein synaptobrevin (yellow) from one membrane forms a tight four-helical bundle with the corresponding SNARE proteins syntaxin-1 (blue) and SNAP25 (red) from a second membrane. The complex brings the membranes close together, initiating the fusion event. [Drawn from 1SFC.pdb.] • The structural details of ligand binding by TLRs (Chapter 34) • The role of the pentose phosphate pathway in rapid cell growth (Chapter 20) • Biochemical characteristics of muscle fiber types (Chapter 21) • Alteration of fatty acid metabolism in tumor cells (Chapter 22) • Biochemical basis of neurological symptoms of phenylketonuria (Chapter 24) • Ribonucleotide reductase as a chemotherapeutic target (Chapter 25) Preface vii • The role of excess choline in the development of heart disease (Chapter 26) • Cycling of the LDL receptor is regulated (Chapter 26) • The role of ceramide metabolism in stimulating tumor growth (Chapter 26) • The extraordinary power of DNA repair systems illustrated by Deinococcus radiodurans (Chapter 28) All of the new media resources for Biochemistry will be available in our new system. www.macmillanhighered.com/launchpad/berg8e LaunchPad is a dynamic, fully integrated learning environment that brings together all of our teaching and learning resources in one place. It also contains the fully interactive e-Book and other newly updated resources for students and instructors, including the following: • NEW Case Studies are a series of biochemistry case studies you can integrate into your course. Each case study gives students practice in working with Figure 34.3 Recognition of a PAMP by a Toll-like receptor. The structure of TLR3 bound to its PAMP, a fragment of double-stranded RNA, as seen from data, developing critical thinking skills, connecting topics, and applying knowledge to real scenarios. We also provide instructional guidance with each case study (with suggestions on how to use the case in the classroom) and aligned assessment questions for quizzes and exams. • Newly Updated Clicker Questions allow instruc tors to integrate active learning in the classroom and to assess students’ understanding of key concepts during lectures. Available in Microsoft Word and PowerPoint students have learned in the book to (PPT). novel medical situ ations. Students read clinical case studies and use basic • Newly Updated Lecture PowerPoints have biochemistry concepts to solve the been developed to minimize preparation time for medical mysteries, applying and new users of the book. These files offer reinforcing what they learn in lecture and suggested lectures from the book. including key illustrations and summaries • Hundreds of self-graded practice prob that instructors can adapt to their teaching lems allow students to test their styles. understanding of concepts explained in the • Updated Layered PPTs deconstruct key text, with immedi ate feedback. concepts, sequences, and processes • The Metabolic Map helps students from the textbook images, allowing under stand the principles and applications instructors to pres ent complex ideas of the core metabolic pathways. Students step-by-step. can work through guided tutorials with • Updated Textbook Images and Tables embedded are offered as high-resolution JPEG files. assessment questions, or explore the Each image has been fully optimized to Metabolic Map on their own using the increase type sizes and adjust color dragging and zooming functionality of the saturation. These images have been map. tested in a large lecture hall to ensure maximum clarity and visibility. • Jmol tutorials by Jeffrey Cohlberg, California the side (top) and from above (bottom). Notice that the PAMP induces • The Clinical Companion, by Gregory receptor dimerization by binding the surfaces on the side of each of the Raner, The University of North Carolina at extracellular domains. [Drawn from 3CIY.pdb]. Greensboro and Douglas Root, University of North Texas, applies concepts that State University at Long Beach, teach students viii how to create models of proteins in Jmol based on data from the Protein Data Bank. By working through the tutorial and answering assessment ques tions at the end of each exercise, students learn to use this important database and fully realize the relationships between the structure and function of enzymes. • Living figures allow students to explore protein structure in 3-D. Students can zoom and rotate the “live” structures to get a better understanding of their three-dimensional nature and can experiment with different display styles (space-filling, ball-and stick, ribbon, backbone) by means of a user-friendly interface. • Concept-based tutorials by Neil D. Clarke help students build an intuitive understanding of some of the more difficult concepts covered in the textbook. • Animated techniques help students grasp experi mental techniques used for exploring genes and proteins. • NEW animations show students biochemical pro cesses in motion. The eighth edition includes many new animations. • Online end-of-chapter questions are assignable and self-graded multiple-choice versions of the end-of-chapter questions in the book, giving stu dents a way to practice applying chapter content in an online environment. • Flashcards are an interactive tool that allows students to study key terms from the book. • LearningCurve is a self-assessment tool that helps students evaluate their progress. Students can test their understanding by taking an online multiple choice quiz provided for each chapter, as well as a general chemistry review. Updated Student Companion [1-4641-8803-3] For each chapter of the textbook, the Student Companion includes: • Chapter Learning Objectives and Summary • Self-Assessment Problems, including multiple choice, short-answer, matching questions, and chal lenge problems, and their answers • Expanded Solutions to end-of-chapter problems in the textbook ix MOLECULAR EVOLUTION This icon signals the start of the many discussions that highlight protein commonalities or other molecular evolutionary insights. Only L amino acids make up proteins (p. 29) Why this set of 20 amino acids? (p. 35) Sickle-cell trait and malaria (p. 206) Additional human globin genes (p. 208) Catalytic triads in hydrolytic enzymes (p. 258) Major classes of peptide-cleaving enzymes (p. 260) Common catalytic core in type II restriction enzymes (p. 275) P-loop NTPase domains (p. 280) Conserved catalytic core in protein kinases (p. 298) Why do different human blood types exist? (p. 331) Archaeal membranes (p. 346) Ion pumps (p. 370) P-type ATPases (p. 374) ATP-binding cassettes (p. 374) Sequence comparisons of Na1 and Ca21 channels (p. 382) Small G proteins (p. 414) Metabolism in the RNA world (p. 444) Why is glucose a prominent fuel? (p. 451) NAD1 binding sites in dehydrogenases (p. 465) Isozymic forms of lactate dehydrogenase (p. 487) Evolution of glycolysis and gluconeogenesis (p. 487) The a-ketoglutarate dehydrogenase complex (p. 505) Domains of succinyl CoA synthetase (p. 507) Evolution of the citric acid cycle (p. 516) Mitochondrial evolution (p. 525) Conserved structure of cytochrome c (p. 541) Common features of ATP synthase and G proteins (p. 548) Pigs lack uncoupling protein 1 (UCP-1) and brown fat (p. 556) Related uncoupling proteins (p. 556) Chloroplast evolution (p. 568) Evolutionary origins of photosynthesis (p. 584) Evolution of the C4 pathway (p. 601) The relationship of the Calvin cycle and the pentose phosphate pathway (p. 610) Increasing sophistication of glycogen phosphorylase regulation (p. 629) Glycogen synthase is homologous to glycogen phosphorylase (p. 631) A recurring motif in the activation of carboxyl groups (p. 649) Prokaryotic counterparts of the ubiquitin pathway and the proteasome (p. 686) A family of pyridoxal-dependent enzymes (p. 692) Evolution of the urea cycle (p. 696) The P-loop NTPase domain in nitrogenase (p. 716) Conserved amino acids in transaminases determine amino acid chirality (p. 721) Feedback inhibition (p. 731) Recurring steps in purine ring synthesis (p. 749) Ribonucleotide reductases (p. 755) Increase in urate levels during primate evolution (p. 761) Deinococcus radiodurans illustrates the power of DNA repair systems (p. 828) DNA polymerases (p. 829) Thymine and the fidelity of the genetic message (p. 849) Sigma factors in bacterial transcription (p. 865) Similarities in transcription between archaea and eukaryotes (p. 876) Evolution of spliceosome-catalyzed splicing (p. 888) Classes of aminoacyl-tRNA synthetases (p. 901) Composition of the primordial ribosome (p. 903) Homologous G proteins (p. 908) A family of proteins with common ligand-binding domains (p. 930) The independent evolution of DNA-binding sites of regulatory proteins (p. 931) Key principles of gene regulation are similar in bacteria and archaea (p. 937) CpG islands (p. 949) Iron-response elements (p. 955) miRNAs in gene evolution (p. 957) The odorant-receptor family (p. 963) Photoreceptor evolution (p. 973) The immunoglobulin fold (p. 988) Relationship of tubulin to prokaryotic proteins (p. 1023) x CLINICAL APPLICATIONS This icon signals the start of a clinical application in the text. Additional, briefer clinical correlations appear in the text as appropriate. Osteogenesis imperfecta (p. 46) Protein-misfolding diseases (p. 56) Protein modification and scurvy (p. 57) Antigen/antibody detection with ELISA (p. 82) Synthetic peptides as drugs (p. 92) PCR in diagnostics and forensics (p.142) Gene therapy (p. 164) Aptamers in biotechnology and medicine (p. 187) Functional magnetic resonance imaging (p. 193) 2,3-BPG and fetal hemoglobin (p. 201) Carbon monoxide poisoning (p. 201) Sickle-cell anemia (p. 205) Thalassemia (p. 207) Aldehyde dehydrogenase deficiency (p. 228) Action of penicillin (p. 239) Protease inhibitors (p. 263) Carbonic anhydrase and osteopetrosis (p. 264) Isozymes as a sign of tissue damage (p. 293) Trypsin inhibitor helps prevent pancreatic damage (p. 302) Emphysema (p. 303) Blood clotting involves a cascade of zymogen activations (p. 303) Vitamin K (p. 306) Antithrombin and hemorrhage (p. 307) Hemophilia (p.308) Monitoring changes in glycosylated hemoglobin (p. 321) Erythropoietin (p. 327) Hurler disease (p. 327) Mucins (p. 329) Blood groups (p. 331) I-cell disease (p. 332) Influenza virus binding (p. 335) Clinical applications of liposomes (p. 349) Aspirin and ibuprofen (p. 353) Digitalis and congestive heart failure (p. 373) Multidrug resistance (p. 374) Long QT syndrome (p. 388) Signal-transduction pathways and cancer (p. 416) Monoclonal antibodies as anticancer drugs (p. 416) Protein kinase inhibitors as anticancer drugs (p. 417) G-proteins, cholera and whooping cough (p. 417) Vitamins (p. 438) Triose phosphate isomerase deficiency (p. 454) Excessive fructose consumption (p. 466) Lactose intolerance (p. 467) Galactosemia (p. 468) Aerobic glycolysis and cancer (p. 474) Phosphatase deficiency (p. 512) Defects in the citric acid cycle and the development of cancer (p. 513) Beriberi and mercury poisoning (p. 515) Frataxin mutations cause Friedreich’s ataxia (p. 531) Reactive oxygen species (ROS) are implicated in a variety of diseases (p. 539) ROS may be important in signal transduction (p. 540) IF1 overexpression and cancer (p. 554) Brown adipose tissue (p. 555) Mild uncouplers sought as drugs (p.557) Mitochondrial diseases (p. 557) Glucose 6-phosphate dehydrogenase deficiency causes drug-induced hemolytic anemia (p. 610) Glucose 6-phosphate dehydrogenase deficiency protects against malaria (p. 612) Developing drugs for type 2 diabetes (p. 636) Glycogen-storage diseases (p. 637) Chanarin-Dorfman syndrome (p. 648) Carnitine deficiency (p. 650) Zellweger syndrome (p. 657) Diabetic ketosis (p. 659) Ketogenic diets to treat epilepsy (p. 660) Some fatty acids may contribute to pathological conditions (p. 661) The use of fatty acid synthase inhibitors as drugs (p. 667) Effects of aspirin on signaling pathways (p. 669) Diseases resulting from defects in transporters of amino acids (p. 682) Diseases resulting from defects in E3 proteins (p. 685) Drugs target the ubiquitin-proteasome system (p.687) Using proteasome inhibitors to treat tuberculosis (p. 687) Blood levels of aminotransferases indicate liver damage (p. 691) Inherited defects of the urea cycle (hyperammonemia) (p. 697) Alcaptonuria, maple syrup urine disease, and phenylketonuria (p. 705) xi High homocysteine levels and vascular disease (p. 726) Inherited disorders of porphyrin metabolism (p. 737) Anticancer drugs that block the synthesis of thymidylate (p. 757) Ribonucleotide reductase is a target for cancer therapy (p. 759) Adenosine deaminase and severe combined immunodefi ciency (p. 760) Gout (p. 761) Lesch–Nyhan syndrome (p. 761) Folic acid and spina bifida (p. 762) Enzyme activation in some cancers to generate phospho choline (p. 770) Excess choline and heart disease (p. 771) Gangliosides and cholera (p. 773) Second messengers derived from sphingolipids and diabetes (p. 773) Respiratory distress syndrome and Tay–Sachs disease (p. 774) Ceramide metabolism stimulates tumor growth (p. 775) Phosphatidic acid phosphatase and lipodystrophy (p. 776) Hypercholesterolemia and atherosclerosis (p. 784) Mutations in the LDL receptor (p. 785) LDL receptor cycling is regulated (p. 787) The role of HDL in protecting against arteriosclerosis (p. 787) Clinical management of cholesterol levels (p. 788) Bile salts are derivatives of cholesterol (p. 789) The cytochrome P450 system is protective (p. 791) A new protease inhibitor also inhibits a cytochrome P450 enzyme (p. 792) Aromatase inhibitors in the treatment of breast and ovarian cancer (p. 794) Rickets and vitamin D (p. 795) Caloric homeostasis is a means of regulating body weight (p. 802) The brain plays a key role in caloric homeostasis (p. 804) Diabetes is a common metabolic disease often resulting from obesity (p. 807) Exercise beneficially alters the biochemistry of cells (p. 813) Food intake and starvation induce metabolic changes (p. 816) Ethanol alters energy metabolism in the liver (p. 819) Antibiotics that target DNA gyrase (p. 839) Blocking telomerase to treat cancer (p. 845) Huntington disease (p. 850) Defective repair of DNA and cancer (p. 850) Detection of carcinogens (Ames test) (p. 852) Translocations can result in diseases (p. 855) Antibiotic inhibitors of transcription (p. 869) Burkitt lymphoma and B-cell leukemia (p. 876) Diseases of defective RNA splicing (p. 884) Vanishing white matter disease (p. 913) Antibiotics that inhibit protein synthesis (p. 914) Diphtheria (p. 914) Ricin, a lethal protein-synthesis inhibitor (p. 915) Induced pluripotent stem cells (p. 947) Anabolic steroids (p. 951) Color blindness (p. 974) The use of capsaicin in pain management (p. 978) Immune-system suppressants (p. 994) MHC and transplantation rejection (p. 1002) AIDS (p. 1003) Autoimmune diseases (p. 1005) Immune system and cancer (p. 1005) Vaccines (p. 1006) Charcot-Marie-Tooth disease (p. 1022) Taxol (p. 1023) xii task. ACKNOWLEDGMENTS Writing a popular textbook is both a challenge and an honor. Our goal is to convey to our students our enthu siasm and understanding of a discipline to which we are devoted. They are our inspiration. Consequently, not a word was written or an illustration constructed with out the knowledge that bright, engaged students would immediately detect vagueness and ambiguity. We also thank our colleagues who supported, advised, instruct ed, and simply bore with us during this arduous Paul Adams University of Arkansas, Fayetteville Kevin Ahern Oregon State University Zulfiqar Ahmad A.T. Still University of Health Sciences Young-Hoon An Wayne State University Richard Amasino University of Wisconsin Kenneth Balazovich University of Michigan Donald Beitz Iowa State University Matthew Berezuk Azusa Pacific University Melanie Berkmen Suffolk University Steven Berry University of Minnesota, Duluth Loren Bertocci Marian University Mrinal Bhattacharjee Long Island University Elizabeth Blinstrup-Good University of Illinois Brian Bothner Montana State University Mark Braiman Syracuse University We are grateful to our colleagues throughout the world who patiently answered our questions and shared their insights into recent developments. We also especially thank those who served as review ers for this new edition. Their thoughtful comments, suggestions, and encouragement have been of immense help to us in maintaining the excellence of the preceding editions. These reviewers are: David Brown Florida Gulf Coast University Donald Burden Middle Tennessee State University Nicholas Burgis Eastern Washington University W. Malcom Byrnes Howard University College of Medicine Graham Carpenter Vanderbilt University School of Medicine John Cogan The Ohio State University Jeffrey Cohlberg California State University, Long Beach David Daleke Indiana University John DeBanzie Northeastern State University Cassidy Dobson St. Cloud State University Donald Doyle Georgia Institute of Technology Ludeman Eng Virginia Tech Caryn Evilia Idaho State University Kirsten Fertuck Northeastern University Brent Feske Armstrong Atlantic University Patricia Flatt Western Oregon University Wilson Francisco Arizona State University Gerald Frenkel Rutgers University Ronald Gary University of Nevada, Las Vegas Eric R. Gauthier Laurentian University Glenda Gillaspy Virginia Tech James Gober UCLA Christina Goode California State University, Fullerton Nina Goodey Montclair State University Eugene Grgory Virginia Tech Robert Grier Atlanta Metropolitan State College Neena Grover Colorado College Paul Hager East Carolina University Ann Hagerman Miami University Mary Hatcher-Skeers Scripps College Diane Hawley University of Oregon Blake Hill Medical College of Wisconsin Pui Ho Colorado State University Charles Hoogstraten Michigan State University Frans Huijing University of Miami Kathryn Huisinga Malone University Cristi Junnes Rocky Mountain College Lori Isom University of Central Arkansas Nitin Jain University of Tennessee Blythe Janowiak Saint Louis University Gerwald Jogl Brown University Kelly Johanson Xavier University of Louisiana Jerry Johnson University of HoustonDowntown Todd Johnson Weber State University David Josephy University of Guelph Michael Kalafatis Cleveland State University Marina Kazakevich University of MassachusettsDartmouth Jong Kim Alabama A&M University xiii Sung-Kun Kim Baylor University Roger Koeppe University of Arkansas, Fayetteville Dmitry Kolpashchikov University of Central Florida Min-Hao Kuo Michigan State University Isabel Larraza North Park University Mark Larson Augustana College Charles Lawrence Montana State University Pan Li State University of New York, Albany Darlene Loprete Rhodes College Greg Marks Carroll University Michael Massiah George Washington University Keri McFarlane Northern Kentucky University Michael Mendenhall University of Kentucky Stephen Mills University of San Diego Smita Mohanty Auburn University Debra Moriarity University of Alabama, Huntsville Stephen Munroe Marquette University Jeffrey Newman Lycoming College William Newton Virginia Tech Alfred Nichols Jacksonville State University Brian Nichols University of Illinois, Chicago Allen Nicholson Temple University Brad Nolen University of Oregon Pamela Osenkowski Loyola University, Chicago Xiaping Pan East Carolina University Stefan Paula Northern Kentucky University David Pendergrass University of KansasEdwards Wendy Pogozelski State University of New York, Geneseo Gary Powell Clemson University Geraldine Prody Western Washington University Joseph Provost University of San Diego Greg Raner University of North Carolina, Greensboro Tanea Reed Eastern Kentucky University Christopher Reid Bryant University Denis Revie California Lutheran University Douglas Root University of North Texas Johannes Rudolph University of Colorado Brian Sato University of California, Irvine Glen Sauer Fairfield University Joel Schildbach Johns Hopkins University Stylianos Scordilis Smith College Ashikh Seethy Maulana Azad Medical College, New Delhi Lisa Shamansky California State University, San Bernardino Bethel Sharma Sewanee: University of the South Nicholas Silvaggi University of Wisconsin-Milwaukee Kerry Smith Clemson University Narashima Sreerama Colorado State University Wesley Stites University of Arkansas Jon Stoltzfus Michigan State University Gerald Stubbs Vanderbilt University Takita Sumter Winthrop University Anna Tan-Wilson State University of New York, Binghamton Steven Theg University of California, Davis Marc Tischler University of Arizona Ken Traxler Bemidji State University Brian Trewyn Colorado School of Mines Vishwa Trivedi Bethune Cookman University Panayiotis Vacratsis University of Windsor Peter van der Geer San Diego State University Jeffrey Voigt Albany College of Pharmacy and Health Sciences Grover Waldrop Louisiana State University Xuemin Wang University of Missouri Yuqi Wang Saint Louis University Rodney Weilbaecher Southern Illinois University Kevin Williams Western Kentucky University Laura Zapanta We have been working with the people at W. H. Freeman/ Macmillan Higher Education for many years now, and our experiences have always been enjoyable and rewarding. Writing and producing the eighth edition of Biochemistry confi rmed our belief that they are a won derful xiv University of Pittsburgh Brent Znosko Saint Louis University publishing team and we are honored to work with them. Our Macmillan colleagues have a knack for undertaking stressful, but exhilarating, projects and reducing the stress without reducing the exhilaration and a remarkable ability to coax without ever nagging. We have many people to thank for this experience, some of whom are fi rst timers to the Biochemistry project. We are delighted to work with Senior Acquisitions Editor, Lauren Schultz, for the fi rst time. She was unfailing in her enthusiasm and generous with her sup port. Another new member of the team was our devel opmental editor, Irene Pech. We have had the pleasure of working with a number of outstanding developmental editors over the years, and Irene continues this tradi tion. Irene is thoughtful, insightful, and very effi cient at identifying aspects of our writing and fi gures that were less than clear. Lisa Samols, a former developmental editor, served as a consultant, archivist for previous edi tions, and a general source of publishing knowledge. Senior Project Editor Deni Showers, with Sherrill Redd, managed the fl ow of the entire project, from copyediting through bound book, with admirable effi ciency. Irene Vartanoff and Mercy Heston, our manuscript editors, enhanced the literary consistency and clarity of the text. Vicki Tomaselli, Design Manager, produced a design and layout that makes the book uniquely attractive while still emphasizing its ties to past editions. Photo Editor Christine Buese and Photo Researcher Jacalyn Wong found the photographs that we hope make the text not only more inviting, but also fun to look through. Janice Donnola, Illustration Coordinator, deftly directed the rendering of new illustrations. Paul Rohloff, Production Coordinator, made sure that the signifi cant diffi culties of scheduling, composition, and manufacturing were smoothly overcome. Amanda Dunning and Donna Brodman did a wonderful job in their management of the media program. In addition, Amanda ably coordi nated the print supplements plan. Special thanks also to editorial assistants Shannon Moloney and Nandini Ahuja. Sandy Lindelof, Executive Marketing Manager, enthusiastically introduced this newest edition of Biochemistry to the academic world. We are deeply appreciative of Craig Bleyer and his sales staff for their support. Without their able and enthusiastic presenta tion of our text to the academic community, all of our efforts would be in vain. We also wish to thank Kate Ahr Parker, Publisher, for her encouragement and belief in us. Thanks also to our many colleagues at our own institu tions as well as throughout the country who patiently answered our questions and encouraged us on our quest. Finally, we owe a debt of gratitude to our fami lies—our wives, Wendie Berg, Alison Unger, and Megan Williams, and our children, especially Timothy and Mark Gatto. Without their support, comfort, and understanding, this endeavor could never have been undertaken, let alone successfully completed. xv BRIEF CONTENTS 3 Exploring Proteins and Proteomes 65 4 DNA, RNA, and the Flow of Genetic Information 105 5 Exploring Genes and Genomes 135 6 Exploring Evolution and Bioinformatics 169 7 Hemoglobin: Portrait of a Protein in Action 191 8 Enzymes: Basic Concepts and Kinetics 215 9 Catalytic Strategies 251 10 Regulatory Strategies 285 11 Carbohydrates 315 12 Lipids and Cell Membranes 341 13 Membrane Channels and Pumps 367 14 Signal-Transduction Pathways 397 Part II TRANSDUCING AND STORING ENERGY 15 Metabolism: Basic Concepts and Design 423 16 Glycolysis and Gluconeogenesis 449 17 The Citric Acid Cycle 495 18 Oxidative Phosphorylation 523 19 The Light Reactions of Photosynthesis 565 20 The Calvin Cycle and the Pentose Phosphate Pathway 589 21 Glycogen Metabolism 617 22 Fatty Acid Metabolism 643 23 Protein Turnover and Amino Acid Catabolism 681 Part III SYNTHESIZING THE MOLECULES OF LIFE 24 The Biosynthesis of Amino Acids 713 25 Nucleotide Biosynthesis 743 26 The Biosynthesis of Membrane Lipids and Steroids 767 27 The Integration of Metabolism 801 28 DNA Replication, Repair, and Recombination 827 29 RNA Synthesis and Processing 859 30 Protein Synthesis 893 31 The Control of Gene Expression in Prokaryotes 925 32 The Control of Gene Expression in Eukaryotes 941 Part IV RESPONDING TO ENVIRONMENTAL CHANGES 33 Sensory Systems 961 34 The Immune System 981 35 Molecular Motors 1011 36 Drug Development 1033 CONTENTS Preface v Part I THE MOLECULAR DESIGN OF LIFE 1 CHAPTER 1 Biochemistry: An Evolving Science 1 CHAPTER 1 Biochemistry: An Evolving Science 1 1.1 Biochemical Unity Underlies Biological Diversity 1 1.2 Part I THE MOLECULAR DESIGN OF LIFE 1 Biochemistry: An Evolving Science 1 2 Protein Composition and Structure 27 DNA Illustrates the Interplay Between Form and Function 4 DNA is constructed from four building blocks 4 Two single strands of DNA combine to form a double helix 5 DNA structure explains heredity and the storage of information 5 1.3 Concepts from Chemistry Explain the Properties of Biological Molecules 6 The formation of the DNA double helix as a key example 6 The double helix can form from its component strands 6 Covalent and noncovalent bonds are CHAPTER 3 Exploring Proteins and Proteomes 65 important for the structure and stability of biological molecules 6 The double helix is an expression of the rules of CHAPTER 3 Exploring Proteins and Proteomes 65 chemistry 9 The laws of thermodynamics govern the behavior of biochemical systems 10 Heat is released in the formation The proteome is the functional representation of of the double helix 12 Acid–base reactions are central in the genome 66 3.1 The Purification of Proteins Is an many biochemical processes 13 Acid–base reactions can Essential First Step in Understanding Their Function 66 disrupt the double helix 14 Buffers regulate pH in organisms The assay: How do we recognize the protein that we are and in the laboratory 15 1.4 The Genomic Revolution Is looking for? 67 Proteins must be released from the cell to Transforming Biochemistry, Medicine, and Other Fields 17 Genome sequencing has transformed biochemistry and other fields 17 Environmental factors influence human biochemistry 20 Genome sequences encode proteins and patterns of expression 21 APPENDIX: Visualizing Molecular Structures I: Small Molecules 22 CHAPTER 2 Protein Composition and Structure 27 CHAPTER 2 Protein Composition and Structure 27 2.1 Proteins Are Built from a Repertoire of 20 Amino Acids 29 2.2 Primary Structure: Amino Acids Are Linked by Peptide be purified 67 Proteins can be purified according to solubility, size, charge, and binding affinity 68 Proteins can be separated by gel electrophoresis and displayed 71 A protein purification scheme can be quantitatively evaluated 75 Ultracentrifugation is valuable for separating biomolecules and determining their masses 76 Protein purification can be made easier with the use of recombinant DNA technology 78 3.2 Immunology Provides Important Techniques with Which to Investigate Proteins 79 Antibodies to specific proteins can be generated 79 Monoclonal antibodies with virtually any desired specificity can be readily prepared 80 Proteins can be detected and quantified by using an enzyme-linked immunosorbent assay 82 Western blotting permits the detection of proteins separated by gel electrophoresis 83 Fluorescent markers make the visualization of proteins in the cell possible 84 Contents xvii Bonds to Form Polypeptide Chains 35 Proteins have unique amino acid sequences specified by genes 37 Polypeptide chains are flexible yet conformationally restricted 38 2.3 Secondary Structure: Polypeptide Chains Can Fold into Regular Structures Such As the Alpha Helix, the Beta Sheet, and Turns and Loops 40 The alpha helix is a coiled structure stabilized by intrachain hydrogen bonds 40 Beta sheets are stabilized by hydrogen bonding between polypeptide strands 42 3.3 Mass Spectrometry Is a Powerful Technique for the Identification of Peptides and Proteins 85 Peptides can be sequenced by mass spectrometry 87 Proteins can be specifically cleaved into small peptides to facilitate analysis 88 Genomic and proteomic methods are complementary 89 The amino acid sequence of a protein provides valuable information 90 Individual proteins can be identified by mass spectrometry 91 3.4 Peptides Can Be Synthesized by Automated Solid-Phase Methods 92 Polypeptide chains can change direction by making reverse turns and loops 44 Fibrous proteins provide structural support for cells and tissues 44 2.4 Tertiary Structure: Water-Soluble Proteins Fold into 3.5 Three-Dimensional Protein Structure Can Be Determined by X-ray Crystallography and NMR Spectroscopy 95 X-ray crystallography reveals three-dimensional structure in atomic detail 95 Nuclear magnetic resonance spectroscopy can reveal the structures of proteins in Compact Structures with Nonpolar Cores 46 2.5 Quaternary Structure: Polypeptide Chains Can Assemble into Multisubunit Structures 48 2.6 The Amino Acid Sequence of a Protein Determines Its Three-Dimensional Structure 49 Amino acids have different propensities for forming a helices, b sheets, and turns 51 Protein folding is a highly cooperative process 52 Proteins fold by progressive stabilization of intermediates rather than by random search 53 Prediction of three-dimensional structure from sequence remains a great challenge 54 Some proteins are inherently unstructured and can exist in multiple conformations 55 Protein misfolding and aggregation are associated with some neurological diseases 56 Protein modification and cleavage confer new capabilities 57 APPENDIX: Visualizing Molecular Structures II: Proteins 61 solution 97 CHAPTER 4 DNA, RNA, and the Flow of CHAPTER 4 DNA, RNA, and the Flow of Genetic Information 105 Genetic Information 105 4.1 A Nucleic Acid Consists of Four Kinds of Bases Linked to a Sugar–Phosphate Backbone 106 RNA and DNA differ in the sugar component and one of the bases 106 Nucleotides are the monomeric units of nucleic acids 107 DNA molecules are very long and have directionality 108 4.2 A Pair of Nucleic Acid Strands with Complementary Sequences Can Form a Double-Helical Structure 109 The double helix is stabilized by genomic DNA 147 Complementary DNA prepared from mRNA can be expressed in host cells 149 Proteins with new functions can be created through directed changes in DNA 150 Recombinant methods enable the exploration of the functional effects of disease-causing mutations 152 hydrogen bonds and van der Waals interactions 109 DNA can assume a variety of structural forms 111 Z-DNA is a 5.3 Complete Genomes Have Been Sequenced and Analyzed left-handed double helix in which backbone phosphates zigzag 112 Some DNA molecules 152 The genomes of organisms ranging from bacteria to are circular and supercoiled 113 Single-stranded nucleic multicellular eukaryotes have been sequenced 153 The sequence of the human genome has been completed 154 acids can adopt elaborate structures 113 4.3 The Double Helix Facilitates the Accurate Transmission of Hereditary Information 114 Differences in DNA density established the validity of the semiconservative replication hypothesis 115 The double helix can be reversibly melted 116 4.4 DNA Is Replicated by Polymerases That Take Instructions from Templates 117 DNA polymerase catalyzes phosphodiester bridge formation 117 The genes of some viruses are made of RNA 118 4.5 Gene Expression Is the Transformation of DNA Information into Functional Molecules 119 Several kinds of RNA play key roles in gene expression 119 xviii Contents All cellular RNA is synthesized by RNA polymerases 120 RNA polymerases take instructions from DNA templates 121 Transcription begins near promoter sites and ends at terminator sites 122 Transfer RNAs are the adaptor molecules in protein synthesis 123 4.6 Amino Acids Are Encoded by Groups of Three Bases Starting from a Fixed Point 124 Major features of the genetic code 125 Messenger RNA contains start and stop signals for protein synthesis 126 The genetic code is nearly universal 126 4.7 Most Eukaryotic Genes Are Mosaics of Introns and Exons 127 RNA processing generates mature RNA 127 Many Next-generation sequencing methods enable the rapid determination of a complete genome sequence 155 Comparative genomics has become a powerful research tool 156 5.4 Eukaryotic Genes Can Be Quantitated and Manipulated with Considerable Precision 157 Geneexpression levels can be comprehensively examined 157 New genes inserted into eukaryotic cells can be efficiently expressed 159 Transgenic animals harbor and express genes introduced into their germ lines 160 Gene disruption and genome editing provide clues to gene function and opportunities for new therapies 160 RNA interference provides an additional tool for disrupting gene expression 162 Tumor-inducing plasmids can be used to introduce new genes into plant cells 163 Human gene therapy holds great promise for medicine 164 CHAPTER 6 Exploring Evolution and CHAPTER 6 Exploring Evolution and Bioinformatics 169 Bioinformatics 169 exons encode protein domains 128 CHAPTER 5 Exploring Genes and Genomes 135 CHAPTER 5 Exploring Genes and Genomes 135 5.1 The Exploration of Genes Relies on Key Tools 136 Restriction enzymes split DNA into specific fragments 137 Restriction fragments can be separated by gel electrophoresis and visualized 137 DNA can be sequenced by controlled termination of replication 138 DNA probes and genes can be synthesized by automated solid-phase methods 139 Selected DNA sequences can be greatly amplified by the polymerase chain reaction 141 PCR is a powerful technique in medical diagnostics, forensics, and studies of molecular evolution 142 The tools for recombinant DNA technology have been used to identify disease-causing mutations 143 5.2 Recombinant DNA Technology Has Revolutionized All Aspects of Biology 143 Restriction enzymes and DNA ligase are key tools in forming recombinant DNA molecules 143 Plasmids and l phage are choice vectors for DNA cloning in bacteria 144 Bacterial and yeast artificial chromosomes 147 Specific genes can be cloned from digests of 6.1 Homologs Are Descended from a Common Ancestor 170 6.2 Statistical Analysis of Sequence Alignments Can Detect Homology 171 The statistical significance of alignments can be estimated by shuffling 173 Distant evolutionary relationships can be detected through the use of substitution matrices 174 Databases can be searched to identify homologous sequences 177 6.3 Examination of Three-Dimensional Structure Enhances Our Understanding of Evolutionary Relationships 177 Tertiary structure is more conserved than primary structure 178 Knowledge of three-dimensional structures can aid in the evaluation of sequence alignments 179 Repeated motifs can be detected by aligning sequences with themselves 180 Convergent evolution illustrates common solutions to biochemical challenges 181 Comparison of RNA sequences can be a source of insight into RNA secondary structures 182 6.4 Evolutionary Trees Can Be Constructed on the Basis of Sequence Information 183 Horizontal gene transfer events may explain unexpected branches of the evolutionary tree 184 6.5 Modern Techniques Make the Experimental Exploration of Evolution Possible 185 Ancient DNA can sometimes be amplified and sequenced 185 Molecular evolution can be examined experimentally 185 Contents xix CHAPTER 7 Hemoglobin: Portrait of a Protein Enzymes: Basic Concepts and CHAPTER 7 Hemoglobin: Portrait of a Protein in Action 191 in Action 191 7.1 Myoglobin and Hemoglobin Bind Oxygen at Iron Atoms in Heme 192 Changes in heme electronic structure upon oxygen binding are the basis for functional imaging studies 193 The structure of myoglobin prevents the release of reactive oxygen species 194 Human hemoglobin is an assembly of four myoglobin like subunits 195 7.2 Hemoglobin Binds Oxygen Cooperatively 195 Oxygen binding markedly changes the quaternary structure of hemoglobin 197 Hemoglobin cooperativity can be potentially explained by several models 198 Structural changes at the heme groups are transmitted to the a1b1– a2b2 interface 200 2,3-Bisphosphoglycerate in red cells is crucial in determining the oxygen affinity of hemoglobin 200 Carbon monoxide can disrupt oxygen transport by hemoglobin 201 7.3 Hydrogen Ions and Carbon Dioxide Promote the Release of Oxygen: The Bohr Effect 202 7.4 Mutations in Genes Encoding Hemoglobin Subunits Can Result in Disease 204 Sickle-cell anemia results from the aggregation of mutated deoxyhemoglobin molecules 205 Thalassemia is caused by an imbalanced production of hemoglobin chains 207 The accumulation of free alpha-hemoglobin chains is prevented 207 Additional globins are encoded in the human genome 208 APPENDIX: Binding Models Can Be Formulated in Quantitative Terms: The Hill Plot and the Concerted Model 210 The formation of an enzyme–substrate complex is the first step in enzymatic catalysis 222 The active sites of enzymes have some common features 223 The binding energy between enzyme and substrate is important for catalysis 225 8.4 The Michaelis–Menten Model Accounts for the Kinetic Properties of Many Enzymes 225 Kinetics is the study of reaction rates 225 The steady-state assumption facilitates a description of enzyme kinetics 226 Variations in KM can have physiological consequences 228 KM and Vmax values can be determined by several means 228 KM and Vmax values are important enzyme characteristics 229 kcat/KM is a measure of catalytic efficiency 230 Most biochemical reactions include multiple substrates 231 Allosteric enzymes do not obey Michaelis– Menten kinetics 233 8.5 Enzymes Can Be Inhibited by Specific Molecules 234 The different types of reversible inhibitors are kinetically distinguishable 235 Irreversible inhibitors can be used to map the active site 237 Penicillin irreversibly inactivates a key enzyme in bacterial cell-wall synthesis 239 Transition-state analogs are potent inhibitors of enzymes 240 Catalytic antibodies demonstrate the importance of selective binding of the transition state to enzymatic activity 241 8.6 Enzymes Can Be Studied One Molecule at a Time 242 APPENDIX: Enzymes are Classified on the Basis of the Types of Reactions That They Catalyze 245 CHAPTER 8 Enzymes: Basic Concepts and CHAPTER 8 Kinetics 215 CHAPTER 9 Catalytic Strategies 251 Formation of the Transition State 221 Kinetics 215 8.1 Enzymes are Powerful and Highly Specific Catalysts 216 Many enzymes require cofactors for activity 217 Enzymes can transform energy from one form into another 217 8.2 Gibbs Free Energy Is a Useful Thermodynamic Function for Understanding Enzymes 218 The free-energy change provides information about the spontaneity but not the rate of a reaction 218 The standard free-energy change of a reaction is related to the equilibrium constant 219 Enzymes alter only the reaction rate and not the reaction equilibrium 220 8.3 Enzymes Accelerate Reactions by Facilitating the CHAPTER 9 Catalytic Strategies 251 A few basic catalytic principles are used by many enzymes 252 9.1 Proteases Facilitate a Fundamentally Difficult Reaction 253 Chymotrypsin possesses a highly reactive serine residue 253 Chymotrypsin action proceeds in two steps linked by a covalently bound intermediate 254 Serine is part of a catalytic triad that also includes histidine and aspartate 255 Catalytic triads are found in other hydrolytic enzymes 258 The catalytic triad has been dissected by site-directed mutagenesis 260 Cysteine, aspartyl, and metalloproteases are other major classes of peptide-cleaving enzymes 260 Protease inhibitors are important drugs 263 xx Contents 9.2 Carbonic Anhydrases Make a Fast Reaction Faster 264 Carbonic anhydrase contains a bound zinc ion essential for catalytic activity 265 Catalysis entails zinc activation of a water molecule 265 A proton shuttle facilitates rapid regeneration of the active form of the enzyme 267 9.3 Restriction Enzymes Catalyze Highly Specific DNA-Cleavage Reactions 269 Cleavage is by in-line displacement of 39-oxygen from phosphorus by magnesium-activated water 269 Restriction enzymes require magnesium for catalytic activity 271 The complete catalytic apparatus is assembled only within complexes of cognate DNA molecules, ensuring specificity 272 Host-cell DNA is protected by the addition of methyl groups to specific bases 274 Type II restriction enzymes have a catalytic core in common and are probably related by horizontal gene transfer 275 9.4 Myosins Harness Changes in Enzyme Conformation to Couple ATP Hydrolysis to Mechanical Work 275 ATP hydrolysis proceeds by the attack of water on the gamma-phosphoryl group 276 Formation of the transition state for ATP hydrolysis is associated with a substantial conformational change 277 The altered conformation of myosin persists for a substantial period of time 278 Scientists can watch single molecules of myosin move 279 Myosins are a family of enzymes containing Cleavage 299 Chymotrypsinogen is activated by specific cleavage of a single peptide bond 299 Proteolytic activation of chymotrypsinogen leads to the formation of a substrate-binding site 300 The generation of trypsin from trypsinogen leads to the activation of other zymogens 301 Some proteolytic enzymes have specific inhibitors 302 Blood clotting is accomplished by a cascade of zymogen activations 303 Prothrombin requires a vitamin Kdependent modification for activation 304 Fibrinogen is converted by thrombin into a fibrin clot 304 Vitamin K is required for the formation of g-carboxyglutamate 306 The clotting process must be precisely regulated 307 Hemophilia revealed an early step in clotting 308 CHAPTER 11 Carbohydrates 315CHAPTER 11 Carbohydrates 315 11.1 Monosaccharides Are the Simplest Carbohydrates 316 Many common sugars exist in cyclic forms 318 Pyranose and furanose rings can assume different conformations 320 Glucose is a reducing sugar 321 Monosaccharides are joined to alcohols and amines through glycosidic bonds 322 Phosphorylated sugars are key intermediates in energy generation and biosyntheses 322 11.2 Monosaccharides Are Linked to Form Complex Carbohydrates 323 Sucrose, lactose, and maltose are P-loop structures 280 CHAPTER 10 Regulatory Strategies 285 CHAPTER 10 Regulatory Strategies 285 10.1 Aspartate Transcarbamoylase Is Allosterically Inhibited by the End Product of Its Pathway 286 Allosterically regulated enzymes do not follow Michaelis–Menten kinetics 287 ATCase consists of separable catalytic and regulatory subunits 287 Allosteric interactions in ATCase are mediated by large changes in quaternary structure 288 Allosteric regulators modulate the T-to-R equilibrium 291 10.2 Isozymes Provide a Means of Regulation Specific to Distinct Tissues and Developmental Stages 292 10.3 Covalent Modification Is a Means of Regulating Enzyme Activity 293 Kinases and phosphatases control the extent of protein phosphorylation 294 Phosphorylation is a highly effective means of regulating the activities of target proteins 296 Cyclic AMP activates protein kinase A by altering the quaternary structure 297 ATP and the target protein bind to a deep cleft in the catalytic subunit of protein kinase A 298 10.4 Many Enzymes Are Activated by Specific Proteolytic the common disaccharides 323 Glycogen and starch are storage forms of glucose 324 Cellulose, a structural component of plants, is made of chains of glucose 324 11.3 Carbohydrates Can Be Linked to Proteins to Form Glycoproteins 325 Carbohydrates can be linked to proteins through asparagine (N-linked) or through serine or threonine (Olinked) residues 326 The glycoprotein erythropoietin is a vital hormone 327 Glycosylation functions in nutrient sensing 327 Proteoglycans, composed of polysaccharides and protein, have important structural roles 327 Proteoglycans are important components of cartilage 328 Mucins are glycoprotein components of mucus 329 Protein glycosylation takes place in the lumen of the endoplasmic reticulum and in the Golgi complex 330 Specific enzymes are responsible for oligosaccharide assembly 331 Blood groups are based on protein glycosylation patterns 331 Errors in glycosylation can result in pathological conditions 332 Oligosaccharides can be “sequenced” 332 11.4 Lectins Are Specific Carbohydrate-Binding Proteins 333 Lectins promote interactions between cells 334 Lectins are organized into different classes 334 Influenza virus binds to sialic acid residues 335 CHAPTER 12 Lipids and Cell Membranes 341 CHAPTER 12 Lipids and Cell Membranes 341 Many common features underlie the diversity of biological membranes 342 12.1 Fatty Acids Are Key Constituents of Lipids 342 Fatty acid names are based on their parent hydrocarbons 342 Fatty acids vary in chain length and degree of unsaturation 343 12.2 There Are Three Common Types of Membrane Lipids 344 Phospholipids are the major class of membrane lipids 344 Membrane lipids can include carbohydrate moieties 345 Cholesterol is a lipid based on a steroid nucleus 346 Archaeal membranes are built from ether lipids with branched chains 346 A membrane lipid is an amphipathic molecule containing a hydrophilic and a hydrophobic moiety 347 12.3 Phospholipids and Glycolipids Readily Form Bimolecular Sheets in Aqueous Media 348 Lipid vesicles can be formed from phospholipids 348 Lipid bilayers are highly impermeable to ions and most polar molecules 349 12.4 Proteins Carry Out Most Membrane Processes 350 Proteins associate with the lipid bilayer in a variety of ways 351 Proteins interact with membranes in a variety of ways 351 Some proteins associate with membranes through covalently attached hydrophobic groups 354 Transmembrane helices can be accurately predicted from amino acid sequences 354 12.5 Lipids and Many Membrane Proteins Diffuse Rapidly in the Plane of the Membrane 356 The fluid mosaic model allows lateral movement but not rotation through the membrane 357 Membrane fluidity is controlled by fatty acid composition and cholesterol content 357 Lipid rafts are highly dynamic complexes formed between cholesterol and specific lipids 358 All biological membranes are asymmetric 358 Contents xxi 12.6 Eukaryotic Cells Contain Compartments Bounded by Internal Membranes 359 CHAPTER 13 Membrane Channels and Pumps 367CHAPTER 13 Membrane Channels and Pumps 367 ATPases couple phosphorylation and conformational changes to pump calcium ions across membranes 370 Digitalis specifically inhibits the Na1–K1 pump by blocking its dephosphorylation 373 P-type ATPases are evolutionarily conserved and play a wide range of roles 374 Multidrug resistance highlights a family of membrane pumps with ATP-binding cassette domains 374 13.3 Lactose Permease Is an Archetype of Secondary Transporters That Use One Concentration Gradient to Power the Formation of Another 376 13.4 Specific Channels Can Rapidly Transport Ions Across Membranes 378 Action potentials are mediated by transient changes in Na1 and K1 permeability 378 Patch-clamp conductance measurements reveal the activities of single channels 379 The structure of a potassium ion channel is an archetype for many ion-channel structures 379 The structure of the potassium ion channel reveals the basis of ion specificity 380 The structure of the potassium ion channel explains its rapid rate of transport 383 Voltage gating requires substantial conformational changes in specific ion-channel domains 383 A channel can be inactivated by occlusion of the pore: the ball-andchain model 384 The acetylcholine receptor is an archetype for ligand-gated ion channels 385 Action potentials integrate the activities of several ion channels working in concert 387 Disruption of ion channels by mutations or chemicals can be potentially life-threatening 388 13.5 Gap Junctions Allow Ions and Small Molecules to Flow Between Communicating Cells 389 13.6 Specific Channels Increase the Permeability of Some Membranes to Water 390 xxii Contents CHAPTER 14 Signal-Transduction Pathways 397 CHAPTER 14 Si l T d i P h Signal-Transduction Pathways 397 Signal transduction depends on molecular circuits 398 14.1 Heterotrimeric G Proteins Transmit Signals and Reset Themselves 399 Ligand binding to 7TM receptors leads to the activation of heterotrimeric G proteins 400 Activated G proteins transmit signals by binding to other proteins 402 Cyclic AMP stimulates the phosphorylation of many target proteins by activating protein kinase A 403 G proteins spontaneously reset themselves through GTP hydrolysis 403 Some 7TM receptors activate the phosphoinositide cascade 404 Calcium ion is a widely used second messenger 405 Calcium ion often activates the regulatory protein calmodulin 407 14.2 Insulin Signaling: Phosphorylation Cascades Are Central to Many Signal-Transduction Processes The expression of transporters largely defines the 407 metabolic activities of a given cell type 368 The insulin receptor is a dimer that closes around a bound 13.1 The Transport of Molecules Across a insulin molecule 408 Insulin binding results in the crossMembrane May Be Active or Passive 368 Many molecules phosphorylation and activation of the insulin receptor 408 require protein transporters to cross membranes 368 Free The activated insulin-receptor kinase initiates a kinase energy stored in concentration gradients can be quantified cascade 409 Insulin signaling is terminated by the action of 369 phosphatases 411 14.3 EGF Signaling: Signal-Transduction 13.2 Two Families of Membrane Proteins Use ATP Hydrolysis Pathways Are Poised to Respond 411 to Pump Ions and Molecules Across Membranes 370 P-type EGF binding results in the dimerization of the EGF receptor 411 The EGF receptor undergoes phosphorylation of its carboxyl-terminal tail 413 EGF signaling leads to the activation of Ras, a small G protein 413 Activated Ras initiates a protein kinase cascade 414 EGF signaling is terminated by protein phosphatases and the intrinsic GTPase activity of Ras 414 14.4 Many Elements Recur with Variation in Different SignalTransduction Pathways 415 14.5 Defects in Signal-Transduction Pathways Can Lead to Cancer and Other Diseases 416 Monoclonal antibodies can be used to inhibit signal transduction pathways activated in tumors 416 Protein kinase inhibitors can be effective anticancer drugs 417 Cholera and whooping cough are the result of altered Gprotein activity 417 Glucose is generated from dietary carbohydrates 450 Glucose is an important fuel for most organisms 451 16.1 Glycolysis Is an Energy-Conversion Pathway in Many Organisms 451 Hexokinase traps glucose in the cell and begins glycolysis 451 Fructose 1,6-bisphosphate is generated from glucose 6-phosphate 453 The six-carbon sugar is cleaved into two three-carbon fragments 454 Mechanism: Triose phosphate isomerase salvages a three-carbon fragment 455 The oxidation of an aldehyde to an acid powers the formation of a compound with high phosphoryl-transfer potential 457 Mechanism: Phosphorylation is coupled to the oxidation of glyceraldehyde 3-phosphate by a thioester intermediate 458 ATP is formed by phosphoryl transfer from 1,3-bisphosphoglycerate 459 Additional ATP is generated with the formation of pyruvate 460 Two ATP molecules are formed in the conversion of glucose into pyruvate 461 Part II TRANSDUCING AND STORING ENERGY NAD1 is regenerated from the metabolism of pyruvate 462 Fermentations provide usable energy in the absence of oxygen 464 The binding site for NAD1 is similar in many dehydrogenases 465 Fructose is converted into glycolytic intermediates by fructokinase 465 Excessive fructose consumption can lead to pathological conditions 466 Galactose is converted into glucose 6-phosphate 466 Many adults are intolerant of milk because they are deficient in lactase 467 Galactose is highly toxic if the transferase is missing 468 CHAPTER 15 Metabolism: Basic Concepts CHAPTER 15 Metabolism: Basic Concepts and Design 423 and Design 423 15.1 Metabolism Is Composed of Many Coupled, Interconnecting Reactions 424 Metabolism consists of energyyielding and energy requiring reactions 424 16.2 The Glycolytic Pathway Is Tightly Controlled 469 Glycolysis in muscle is regulated to meet the need for ATP A thermodynamically unfavorable reaction can be driven 469 The regulation of glycolysis in the liver illustrates the by a favorable reaction 425 biochemical versatility of the liver 472 A family of transporters enables glucose to enter and leave animal 15.2 ATP Is the Universal Currency of Free cells 473 Aerobic glycolysis is a property of rapidly growing Energy in Biological Systems 426 ATP hydrolysis is exergonic cells 474 Cancer and endurance training affect glycolysis in 426 ATP hydrolysis drives metabolism by shifting the a similar fashion 476 equilibrium of coupled reactions 427 The high phosphoryl potential of ATP results from structural differences 16.3 Glucose Can Be Synthesized from between ATP and its hydrolysis products 429 Phosphoryl- Noncarbohydrate Precursors 476 Gluconeogenesis is not a transfer potential is an important form of cellular energy reversal of glycolysis 478 The conversion of pyruvate into transformation 430 phosphoenolpyruvate begins with the formation of 15.3 The Oxidation of Carbon Fuels Is an Important Source of Cellular Energy 432 Compounds with high phosphoryl-transfer potential can couple carbon oxidation to ATP synthesis 432 Ion gradients across membranes provide an important form of cellular energy that can be coupled to ATP synthesis 433 Phosphates play a prominent role in biochemical processes 434 Energy from foodstuffs is extracted in three stages 434 oxaloacetate 478 Oxaloacetate is shuttled into the cytoplasm and converted into phosphoenolpyruvate 480 The conversion of fructose 1,6-bisphosphate into fructose 6-phosphate and orthophosphate is an irreversible step 480 The generation of free glucose is an important control point 481 Six high-transfer-potential phosphoryl groups are spent in synthesizing glucose from pyruvate 481 16.4 Gluconeogenesis and Glycolysis Are 15.4 Metabolic Pathways Contain Many Recurring Motifs 435 Activated carriers exemplify the modular Reciprocally Regulated 482 Energy charge determines design and economy of metabolism 435 Many activated whether glycolysis or carriers are derived from vitamins 438 Key reactions are gluconeogenesis will be most active 482 The balance between glycolysis and gluconeogenesis in the liver is reiterated throughout metabolism 440 Metabolic processes sensitive to blood-glucose concentration 483 Substrate are regulated in three principal ways 442 Aspects of cycles amplify metabolic signals and metabolism may have evolved from an RNA world 444 CHAPTER 16 Glycolysis and produce heat 485 Lactate and alanine formed by contracting muscle are used by other organs 485 Glycolysis and gluconeogenesis are evolutionarily intertwined 487 Gluconeogenesis 449CHAPTER Gluconeogenesis 449 16 Glycolysis and CHAPTER 17 The Citric Acid Cycle 495 CHAPTER 17 The xxiv Contents Citric Acid Cycle 495 The citric acid cycle harvests high-energy electrons 496 17.1 The Pyruvate Dehydrogenase Complex Links Glycolysis to the Citric Acid Cycle 497 Mechanism: The synthesis of acetyl coenzyme A from pyruvate requires three enzymes and five coenzymes 498 Contents xxiii Flexible linkages allow lipoamide to move between different active sites 500 17.2 The Citric Acid Cycle Oxidizes Two-Carbon Units 501 Citrate synthase forms citrate from The high-potential electrons of NADH enter the respiratory chain at NADH-Q oxidoreductase 532 Ubiquinol is the entry point for electrons from FADH2 of flavoproteins 533 Electrons flow from ubiquinol to cytochrome c through Q-cytochrome c oxidoreductase 533 The Q cycle funnels electrons from a two-electron carrier to a one-electron carrier and pumps protons 535 Cytochrome c oxidase catalyzes the reduction of molecular oxygen to water 535 Toxic derivatives of molecular oxygen such as superoxide radicals are scavenged by protective enzymes 538 Electrons can be transferred between groups that are not in contact 540 The conformation of cytochrome c has remained essentially constant for more than a billion years 541 18.4 A oxaloacetate and acetyl coenzyme A 502 Mechanism: The Proton Gradient Powers the Synthesis of ATP 541 mechanism of citrate synthase prevents undesirable ATP synthase is composed of a proton-conducting unit and reactions 502 Citrate is isomerized into isocitrate 504 a catalytic unit 543 Proton flow through ATP synthase leads Isocitrate is oxidized and decarboxylated to alpha to the release of tightly bound ATP: The binding-change ketoglutarate 504 Succinyl coenzyme A is formed by the mechanism 544 Rotational catalysis is the world’s smallest oxidative molecular motor 546 Proton flow around the c ring powers decarboxylation of alpha-ketoglutarate 505 A compound ATP synthesis 546 ATP synthase and G proteins have with high phosphoryl-transfer potential is generated from several common features 548 succinyl coenzyme A 505 Mechanism: Succinyl coenzyme 18.5 Many Shuttles Allow Movement Across Mitochondrial A synthetase transforms types of biochemical energy 506 Oxaloacetate is regenerated by the oxidation of succinate Membranes 549 Electrons from cytoplasmic NADH enter mitochondria by shuttles 549 The entry of ADP into 507 The citric acid cycle produces high-transfer-potential mitochondria is coupled to the exit of ATP by ATP-ADP electrons, ATP, and CO2 508 translocase 550 Mitochondrial transporters for metabolites 17.3 Entry to the Citric Acid Cycle and Metabolism Through It have a Are Controlled 510 The pyruvate dehydrogenase complex is common tripartite structure 551 18.6 The Regulation of regulated Cellular Respiration Is Governed Primarily by the Need for ATP allosterically and by reversible phosphorylation 511 The 552 The complete oxidation of glucose yields about citric acid cycle is controlled at several points 512 Defects in 30 molecules of ATP 552 The rate of oxidative the citric acid cycle contribute to the phosphorylation is determined by the need for ATP 553 development of cancer 513 ATP synthase can be regulated 554 Regulated uncoupling 17.4 The Citric Acid Cycle Is a Source of Biosynthetic Precursors 514 The citric acid cycle must be capable of being rapidly replenished 514 The disruption of pyruvate metabolism is the cause of beriberi and poisoning by mercury and arsenic 515 The citric acid cycle may have evolved from preexisting pathways 516 leads to the generation of heat 554 Oxidative phosphorylation can be inhibited at many stages 556 Mitochondrial diseases are being discovered 557 Mitochondria play a key role in apoptosis 557 Power transmission by proton gradients is a central motif of 17.5 The Glyoxylate Cycle Enables Plants and Bacteria to Grow bioenergetics 558 on Acetate 516 CHAPTER 19 The Light Reactions of CHAPTER 18 Oxidative Phosphorylation 523CHAPTER 18 Oxidative Phosphorylation 523 18.1 Eukaryotic Oxidative Phosphorylation Takes Place in Mitochondria 524 Mitochondria are bounded by a double membrane 524 Mitochondria are the result of an endosymbiotic event 525 18.2 Oxidative Phosphorylation Depends on Electron Transfer 526 The electron-transfer potential of an electron is measured as redox potential 526 A 1.14-volt potential difference between NADH and molecular oxygen drives electron transport through the chain and favors the formation of a proton gradient 528 18.3 The Respiratory Chain Consists of Four Complexes: Three Proton Pumps and a Physical Link to the Citric Acid Cycle 529 Iron–sulfur clusters are common components of the electron transport chain 531 CHAPTER 19 The Light Reactions of Photosynthesis 565 Photosynthesis 565 Photosynthesis converts light energy into chemical energy 566 19.1 Photosynthesis Takes Place in Chloroplasts 567 The primary events of photosynthesis take place in thylakoid membranes 567 Chloroplasts arose from an endosymbiotic event 568 19.2 Light Absorption by Chlorophyll Induces Electron Transfer 568 A special pair of chlorophylls initiate charge separation 569 Cyclic electron flow reduces the cytochrome of the reaction center 572 19.3 Two Photosystems Generate a Proton Gradient and Crassulacean acid metabolism permits growth in arid ecosystems 601 NADPH in Oxygenic Photosynthesis 572 Photosystem II transfers electrons from water to plastoquinone and generates a proton gradient 572 Cytochrome bf links photosystem II to photosystem I 575 Photosystem I uses light energy to generate reduced ferredoxin, a powerful reductant 575 Ferredoxin–NADP1 reductase converts NADP1 into NADPH 576 20.3 The Pentose Phosphate Pathway Generates NADPH and Synthesizes Five-Carbon Sugars 601 Two molecules of NADPH are generated in the conversion of glucose 6-phosphate into ribulose 5-phosphate 602 The pentose phosphate pathway and glycolysis are linked by transketolase and transaldolase 19.4 A Proton Gradient across the Thylakoid Membrane Drives 602 Mechanism: Transketolase and transaldolase stabilize ATP Synthesis 578 The ATP synthase of chloroplasts closely carbanionic intermediates by different mechanisms 605 resembles those of mitochondria and prokaryotes 578 The 20.4 The Metabolism of Glucose 6-Phosphate by the Pentose activity of chloroplast ATP synthase is regulated 579 Cyclic electron flow through photosystem I leads to the production of Phosphate Pathway Is Coordinated with Glycolysis 607 ATP instead of NADPH 580 The absorption of eight photons The rate of the pentose phosphate pathway is controlled yields one O2, two NADPH, and three ATP molecules 581 by the level of NADP1 607 The flow of glucose 6-phosphate depends on the need for NADPH, ribose 5-phosphate, and 19.5 Accessory Pigments Funnel Energy into Reaction Centers ATP 608 The pentose phosphate pathway is required for 581 Resonance energy transfer allows energy to move from rapid cell growth 610 Through the looking-glass: The the site of initial absorbance to the reaction center 582 The Calvin cycle and the pentose phosphate pathway are components of photosynthesis are highly organized 583 Many mirror images 610 herbicides inhibit the light reactions of 20.5 Glucose 6-Phosphate Dehydrogenase Plays a Key Role in photosynthesis 584 19.6 The Ability to Convert Light into Protection Against Reactive Oxygen Species 610 Glucose 6Chemical Energy Is Ancient 584 phosphate dehydrogenase deficiency Artificial photosynthetic systems may provide clean, causes a drug-induced hemolytic anemia 610 A deficiency of glucose 6-phosphate dehydrogenase confers an evolutionary advantage in some circumstances 612 Contents xxv renewable energy 585 Muscle phosphorylase is regulated by the intracellular energy charge 625 Biochemical characteristics of muscle fiber types differ 625 Phosphorylation promotes the conversion of phosphorylase b to phosphorylase a 626 Phosphorylase kinase is activated by phosphorylation and calcium ions 626 CHAPTER 20 The Calvin Cycle and the CHAPTER 20 The Calvin Cycle and the Pentose Phosphate Pathway 589 Pentose Phosphate Pathway 589 20.1 The Calvin Cycle Synthesizes Hexoses from Carbon Dioxide and Water 590 Carbon dioxide reacts with ribulose 1,5bisphosphate to form two molecules of 3-phosphoglycerate 591 Rubisco activity depends on magnesium and carbamate 592 Rubisco activase is essential for rubisco activity 593 Rubisco also catalyzes a wasteful oxygenase reaction: Catalytic imperfection 593 Hexose phosphates are made from phosphoglycerate, and ribulose 1,5-bisphosphate is regenerated 594 Three ATP and two NADPH molecules are used to bring carbon dioxide to the level of a hexose 597 Starch and sucrose are the major carbohydrate stores in plants 597 20.2 The Activity of the Calvin Cycle Depends on Environmental Conditions 598 Rubisco is activated by light-driven changes in proton and magnesium ion concentrations 598 Thioredoxin plays a key role in regulating the Calvin cycle 599 The C4 pathway of tropical plants accelerates photosynthesis by concentrating carbon dioxide 599 CHAPTER 21 Glycogen Metabolism 617 21.3 Epinephrine and Glucagon Signal the Need for Glycogen Breakdown 627 G proteins transmit the signal for the initiation of glycogen breakdown 627 Glycogen breakdown must be rapidly turned off when necessary 629 The regulation of glycogen phosphorylase became more sophisticated as the enzyme evolved 629 21.4 Glycogen Is Synthesized and Degraded by Different Pathways 630 UDP-glucose is an activated form of glucose 630 Glycogen synthase catalyzes the transfer of glucose from UDP-glucose to a growing chain 630 A branching enzyme forms a-1,6 linkages 631 Glycogen synthase is the key regulatory enzyme in glycogen synthesis 632 Glycogen is an efficient storage form of glucose 632 21.5 Glycogen Breakdown and Synthesis Are Reciprocally Regulated 632 Protein phosphatase 1 reverses the regulatory effects of kinases on glycogen metabolism 633 Insulin stimulates glycogen synthesis by inactivating glycogen synthase kinase 635 Glycogen metabolism in the liver regulates the blood-glucose level 635 A biochemical understanding of glycogen-storage diseases is possible 637 CHAPTER 22 Fatty Acid Metabolism Glycogen metabolism is the regulated release and storage of glucose 618 Glycogen CHAPTER 21 Metabolism 617 21.1 Glycogen Breakdown Requires the Interplay of Several Enzymes 619 Phosphorylase catalyzes the phosphorolytic cleavage of glycogen to release glucose 1-phosphate 619 643 Mechanism: Pyridoxal phosphate participates in the phosphorolytic cleavage of glycogen 620 A debranching enzyme also is needed for the breakdown of glycogen 621 Phosphoglucomutase converts glucose 1-phosphate into glucose 6-phosphate 622 The liver contains glucose 6phosphatase, a hydrolytic enzyme absent from muscle 622 21.2 Phosphorylase Is Regulated by Allosteric Interactions and Reversible Phosphorylation 623 Liver phosphorylase produces glucose for use by other tissues 623 CHAPTER 22 22.6 Acetyl CoA Carboxylase Plays a Key Role in Controlling Fatty Acid Metabolism 670 Acetyl CoA carboxylase is regulated by conditions in the cell 671 Acetyl CoA carboxylase is regulated by a variety of hormones 671 CHAPTER 23 Protein Turnover and CHAPTER 23 Protein Turnover and Fatty Acid Metabolism Amino Acid Catabolism 681 Amino Acid Catabolism 681 643 23.1 Proteins are Degraded to Amino Acids 682 The digestion of Fatty acid degradation and synthesis mirror each other in their chemical reactions 644 22.1 Triacylglycerols Are Highly Concentrated Energy Stores 645 Dietary lipids are digested by pancreatic lipases 645 Dietary lipids are transported in chylomicrons 646 22.2 The Use of Fatty Acids as Fuel Requires Three Stages of Processing 647 Triacylglycerols are hydrolyzed by hormone stimulated lipases 647 Free fatty acids and glycerol are released into the blood 648 Fatty acids are linked to coenzyme A before they are oxidized 648 Carnitine carries long-chain activated fatty acids into the mitochondrial matrix 649 Acetyl CoA, NADH, and FADH2 are generated in each round of fatty acid oxidation 650 xxvi Contents The complete oxidation of palmitate yields 106 molecules of ATP 652 22.3 Unsaturated and Odd-Chain Fatty Acids Require Additional Steps for Degradation 652 An isomerase and a reductase are required for the oxidation of unsaturated fatty acids 652 Odd-chain fatty acids yield propionyl CoA in the final thiolysis step 654 Vitamin B12 contains a corrin ring and a cobalt atom 654 Mechanism: Methylmalonyl CoA mutase catalyzes a rearrangement to form succinyl CoA 655 Fatty acids are also oxidized in peroxisomes 656 Ketone bodies are formed from acetyl CoA when fat breakdown predominates 657 Ketone bodies are a major fuel in some tissues 658 Animals cannot convert fatty acids into glucose 660 Some fatty acids may contribute to the development of pathological conditions 661 22.4 Fatty Acids Are Synthesized by Fatty Acid Synthase 661 Fatty acids are synthesized and degraded by different pathways 661 The formation of malonyl CoA is the committed step in fatty acid synthesis 662 Intermediates in fatty acid synthesis are attached to an acyl carrier protein 662 Fatty acid synthesis consists of a series of condensation, reduction, dehydration, and reduction reactions 662 Fatty acids are synthesized by a multifunctional enzyme complex in animals 664 The synthesis of palmitate requires 8 molecules of acetyl CoA, 14 molecules of NADPH, and 7 molecules of ATP 666 Citrate carries acetyl groups from mitochondria to the cytoplasm for fatty acid synthesis 666 Several sources supply NADPH for fatty acid synthesis 667 Fatty acid metabolism is altered in tumor cells 667 22.5 The Elongation and Unsaturation of Fatty Acids are Accomplished by Accessory Enzyme Systems 668 Membrane-bound enzymes generate unsaturated fatty acids 668 Eicosanoid hormones are derived from polyunsaturated fatty acids 669 Variations on a theme: Polyketide and nonribosomal peptide synthetases resemble fatty acid synthase 670 dietary proteins begins in the stomach and is completed in the intestine 682 Cellular proteins are degraded at different rates 682 23.2 Protein Turnover Is Tightly Regulated 683 Ubiquitin tags proteins for destruction 683 The proteasome digests the ubiquitin-tagged proteins 685 The ubiquitin pathway and the proteasome have prokaryotic counterparts 686 Protein degradation can be used to regulate biological function 687 23.3 The First Step in Amino Acid Degradation Is the Removal of Nitrogen 687 Alpha-amino groups are converted into ammonium ions by the oxidative deamination of glutamate 687 Mechanism: Pyridoxal phosphate forms Schiff-base intermediates in aminotransferases 689 Aspartate aminotransferase is an archetypal pyridoxal dependent transaminase 690 Blood levels of aminotransferases serve a diagnostic function 691 Pyridoxal phosphate enzymes catalyze a wide array of reactions 691 Serine and threonine can be directly deaminated 692 Peripheral tissues transport nitrogen to the liver 692 23.4 Ammonium Ion Is Converted into Urea in Most Terrestrial Vertebrates 693 The urea cycle begins with the formation of carbamoyl phosphate 693 Carbamoyl phosphate synthetase is the key regulatory enzyme for urea synthesis 694 Carbamoyl phosphate reacts with ornithine to begin the urea cycle 694 The urea cycle is linked to gluconeogenesis 696 Urea-cycle enzymes are evolutionarily related to enzymes in other metabolic pathways 696 Inherited defects of the urea cycle cause hyperammonemia and can lead to brain damage 697 Urea is not the only means of disposing of excess nitrogen 698 23.5 Carbon Atoms of Degraded Amino Acids Emerge as Major Metabolic Intermediates 698 Pyruvate is an entry point into metabolism for a number of amino acids 699 Oxaloacetate is an entry point into metabolism for aspartate and asparagine 700 Alphaketoglutarate is an entry point into metabolism for fivecarbon amino acids 700 Succinyl coenzyme A is a point of entry for several nonpolar amino acids 701 Methionine degradation requires the formation of a key methyl donor, S-adenosylmethionine 701 The branched-chain amino acids yield acetyl CoA, acetoacetate, or propionyl CoA 701 Oxygenases are required for the degradation of aromatic amino acids 703 23.6 Inborn Errors of Metabolism Can Disrupt Amino Acid Degradation 705 Phenylketonuria is one of the most common metabolic disorders 706 Determining the basis of the neurological symptoms of phenylketonuria is an active area of research 706 Part III SYNTHESIZING THE MOLECULES OF LIFE CHAPTER 24 The Biosynthesis of Amino Acids 713 CHAPTER 24 The Biosynthesis of Amino Acids 713 Amino acid synthesis requires solutions to three key biochemical problems 714 24.1 Nitrogen Fixation: Microorganisms Use ATP and a Powerful Reductant to Reduce Atmospheric Nitrogen to Ammonia 714 The iron–molybdenum cofactor of nitrogenase binds and reduces atmospheric nitrogen 715 Ammonium ion is assimilated into an amino acid through glutamate and glutamine 717 24.2 Amino Acids Are Made from Intermediates of the Citric Acid Cycle and Other Major Pathways 719 Human beings can synthesize some amino acids but must obtain others from their diet 719 Aspartate, alanine, and glutamate are formed by the addition of an amino group to an alpha-ketoacid 720 A common step determines the chirality of all amino acids 721 The formation of asparagine from aspartate requires an adenylated intermediate 721 Glutamate is the precursor of glutamine, proline, and arginine 722 3-Phosphoglycerate is the precursor of serine, cysteine, and glycine 722 Tetrahydrofolate carries activated one-carbon units at several oxidation levels 723 S-Adenosylmethionine is the major donor of methyl groups 724 Cysteine is synthesized from serine and homocysteine 726 High homocysteine levels correlate with vascular disease 726 Shikimate and chorismate are intermediates in the biosynthesis of aromatic amino acids 727 Tryptophan synthase illustrates substrate channeling in enzymatic catalysis 729 24.3 Feedback Inhibition Regulates Amino Acid Biosynthesis 730 Branched pathways require sophisticated regulation 731 The sensitivity of glutamine synthetase to allosteric regulation is altered by covalent modification 732 24.4 Amino Acids Are Precursors of Many Biomolecules 734 Glutathione, a gamma-glutamyl peptide, serves as a sulfhydryl buffer and an antioxidant 734 Nitric oxide, a short-lived signal molecule, is formed from arginine 735 Porphyrins are synthesized from glycine and succinyl coenzyme A 736 Porphyrins accumulate in some a pyrimidine nucleotide and is converted into uridylate 746 Nucleotide mono-, di-, and triphosphates are interconvertible 747 CTP is formed by amination of UTP 747 Salvage pathways recycle pyrimidine bases 748 25.2 Purine Bases Can Be Synthesized de Novo or Recycled by Salvage Pathways 748 The purine ring system is assembled on ribose phosphate 749 The purine ring is assembled by successive steps of activation by phosphorylation followed by displacement 749 AMP and GMP are formed from IMP 751 Enzymes of the purine synthesis pathway associate with one another in vivo 752 Salvage pathways economize intracellular energy expenditure 752 25.3 Deoxyribonucleotides Are Synthesized by the Reduction of Ribonucleotides Through a Radical Mechanism 753 Mechanism: A tyrosyl radical is critical to the action of ribonucleotide reductase 753 Stable radicals other than tyrosyl radical are employed by other ribonucleotide reductases 755 Thymidylate is formed by the methylation of deoxyuridylate 755 Dihydrofolate reductase catalyzes the regeneration of tetrahydrofolate, a one-carbon carrier 756 Several valuable anticancer drugs block the synthesis of thymidylate 757 25.4 Key Steps in Nucleotide Biosynthesis Are Regulated by Feedback Inhibition 758 Pyrimidine biosynthesis is regulated by aspartate transcarbamoylase 758 The synthesis of purine nucleotides is controlled by feedback inhibition at several sites 758 The synthesis of deoxyribonucleotides is controlled by the regulation of ribonucleotide reductase 759 25.5 Disruptions in Nucleotide Metabolism Can Cause Pathological Conditions 760 The loss of adenosine deaminase activity results in severe combined immunodeficiency 760 Gout is induced by high serum levels of urate 761 Lesch–Nyhan syndrome is a dramatic consequence of mutations in a salvagepathway enzyme 761 Folic acid deficiency promotes birth defects such as spina bifida 762 xxviii Contents CHAPTER 26 The Biosynthesis of Membrane CHAPTER 26 The Biosynthesis of Membrane Lipids and inherited disorders of porphyrin metabolism 737 CHAPTER 25 Nucleotide Biosynthesis 743 CHAPTER 25 Steroids 767 Lipids and Steroids 767 Nucleotide Biosynthesis 743 Nucleotides can be synthesized by de novo or salvage pathways 744 26.1 Phosphatidate Is a Common Intermediate in the Contents xxvii Synthesis of Phospholipids and Triacylglycerols 768 The synthesis of phospholipids requires 25.1 The Pyrimidine Ring Is Assembled de Novo or Recovered by Salvage Pathways 744 Bicarbonate and other oxygenated carbon compounds are activated by phosphorylation 745 The side chain of glutamine can be hydrolyzed to generate ammonia 745 Intermediates can move between active sites by channeling 745 Orotate acquires a ribose ring from PRPP to form an activated intermediate 769 Some phospholipids are synthesized from an activated alcohol 770 Phosphatidylcholine is an abundant phospholipid 770 Excess choline is implicated in the development of heart disease 771 Base-exchange reactions can generate phospholipids 771 Sphingolipids are synthesized from ceramide 772 Gangliosides are carbohydrate-rich sphingolipids that contain acidic sugars 772 Sphingolipids confer diversity on lipid structure and function 773 Respiratory distress syndrome and Tay–Sachs disease result from the disruption of lipid metabolism 774 Ceramide metabolism stimulates tumor growth 774 Phosphatidic acid phosphatase is a key regulatory enzyme in lipid metabolism 775 26.2 Cholesterol Is Synthesized from Acetyl Coenzyme A in Three Stages 776 The synthesis of mevalonate, which is activated as isopentenyl pyrophosphate, initiates the synthesis of cholesterol 776 Squalene (C30) is synthesized from six molecules of isopentenyl pyrophosphate (C5) 777 Squalene cyclizes to form cholesterol 778 26.3 The Complex Regulation of Cholesterol Biosynthesis Takes Place at Several Levels 779 Lipoproteins transport cholesterol and triacylglycerols throughout the organism 782 Low-density lipoproteins play a central role in cholesterol metabolism 784 The absence of the LDL receptor leads to hypercholesterolemia and atherosclerosis 784 Mutations in the LDL receptor prevent LDL release and result in receptor destruction 785 Cycling of the LDL receptor is regulated 787 HDL appears to protect against atherosclerosis 787 The clinical management of cholesterol levels can be understood at a biochemical level 788 26.4 Important Derivatives of Cholesterol Include Bile Salts and Steroid Hormones 788 Letters identify the steroid rings and numbers identify the carbon atoms 790 Steroids are hydroxylated by cytochrome P450 monooxygenases that use NADPH and O2 790 The cytochrome P450 system is widespread and performs a protective function 791 Pregnenolone, a precursor of many other steroids, is formed from cholesterol by cleavage of its side chain 792 Progesterone and corticosteroids are synthesized from pregnenolone 792 Androgens and estrogens are synthesized from pregnenolone 792 Vitamin D is derived from cholesterol by the ring Beneficially Alters the Biochemistry of Cells 813 Mitochondrial biogenesis is stimulated by muscular activity 813 Fuel choice during exercise is determined by the intensity and duration of activity 813 27.5 Food Intake and Starvation Induce Metabolic Changes 816 The starved–fed cycle is the physiological response to a fast 816 Metabolic adaptations in prolonged starvation minimize protein degradation 818 27.6 Ethanol Alters Energy Metabolism in the Liver 819 Ethanol metabolism leads to an excess of NADH 820 Excess ethanol consumption disrupts vitamin metabolism 821 CHAPTER 28 DNA Replication, Repair, CHAPTER 28 DNA Replication, Repair, and Recombination 827 and Recombination 827 28.1 DNA Replication Proceeds by the Polymerization of Deoxyribonucleoside Triphosphates Along a Template 828 DNA polymerases require a template and a primer 829 All DNA polymerases have structural features in common 829 Two bound metal ions participate in the polymerase reaction 829 The specificity of replication is dictated by complementarity of shape between bases 830 An RNA primer synthesized by primase enables DNA synthesis to begin 831 One strand of DNA is made continuously, whereas the other strand is synthesized in fragments 831 DNA ligase joins ends of DNA in duplex regions 832 The separation of DNA strands requires specific helicases and ATP hydrolysis 832 28.2 DNA Unwinding and Supercoiling Are Controlled by splitting activity of light 794 Topoisomerases 833 The linking number of DNA, a topological property, determines the degree of supercoiling 835 CHAPTER 27 The Integration of Metabolism 801 CHAPTER Topoisomerases prepare the double helix for unwinding 836 Type I topoisomerases relax supercoiled 27 The Integration of Metabolism 801 structures 836 Type II topoisomerases can introduce negative supercoils through coupling to ATP hydrolysis 837 28.3 27.1 Caloric Homeostasis Is a Means of Regulating Body Weight 802 27.2 The Brain Plays a Key Role in Caloric DNA Replication Is Highly Coordinated 839 DNA replication Homeostasis 804 Signals from the gastrointestinal tract induce requires highly processive feelings of satiety 804 Leptin and insulin regulate longpolymerases 839 The leading and lagging strands are term control over caloric homeostasis 805 Leptin is one of synthesized in a coordinated fashion 840 DNA replication several hormones secreted by in Escherichia coli begins at a adipose tissue 806 Leptin resistance may be a contributing unique site 842 DNA synthesis in eukaryotes is initiated at factor to multiple sites 843 Telomeres are unique structures at the obesity 806 Dieting is used to combat obesity 807 27.3 ends of linear chromosomes 844 Telomeres are replicated Diabetes Is a Common Metabolic Disease Often Resulting from by telomerase, a specialized polymerase that carries its own RNA template 845 Obesity 807 Insulin initiates a complex signal-transduction pathway in muscle 808 Metabolic syndrome often precedes 28.4 Many Types of DNA Damage Can Be type 2 diabetes 809 Excess fatty acids in muscle modify Repaired 845 Errors can arise in DNA replication 846 Bases metabolism 810 Insulin resistance in muscle facilitates can be damaged by oxidizing agents, alkylating agents, pancreatic failure 810 Metabolic derangements in type 1 and light 846 DNA damage can be detected and repaired by a variety of systems 847 The presence of thymine diabetes result from instead of uracil in DNA permits the repair of deaminated insulin insufficiency and glucagon excess 812 27.4 Exercise cytosine 849 Some genetic diseases are caused by the expansion of repeats of three nucleotides 850 Many cancers are caused by the defective repair of DNA 850 Many potential carcinogens can be detected by their mutagenic action on bacteria 852 28.5 DNA Recombination Plays Important Roles in Replication, Repair, and Other Processes 852 RecA can initiate recombination by promoting strand invasion 853 Some recombination reactions proceed through Holliday-junction intermediates 854 to Both Mechanism and Evolution 886 xxx Contents CHAPTER 30 Protein Synthesis 893 CHAPTER 30 P iS hrotein Synthesis 893 30.1 Protein Synthesis Requires the Translation of Nucleotide Sequences into Amino Acid Sequences 894 The synthesis of Contents xxix long proteins requires a low error frequency 894 Transfer RNA molecules have a common design 895 Some transfer RNA molecules recognize more than one codon because of wobble in base-pairing 897 CHAPTER 29 RNA S h i d P RNA Synthesis and Processing 30.2 Aminoacyl Transfer RNA Synthetases Read the Genetic Code 898 Amino acids are first activated by adenylation 898 Aminoacyl-tRNA synthetases have highly RNA synthesis comprises three stages: Initiation, discriminating amino acid activation sites 899 Proofreading elongation, and termination 860 29.1 RNA Polymerases by aminoacyl-tRNA synthetases Catalyze Transcription 861 RNA chains are formed de novo increases the fidelity of protein synthesis 900 Synthetases and grow in the recognize various features of transfer 59-to-39 direction 862 RNA polymerases backtrack and RNA molecules 901 Aminoacyl-tRNA synthetases can be divided into two classes 901 correct errors 863 RNA polymerase binds to promoter sites on the 30.3 The Ribosome Is the Site of Protein Synthesis 902 DNA template to initiate transcription 864 Sigma subunits Ribosomal RNAs (5S, 16S, and 23S rRNA) play a central of RNA polymerase recognize role in protein synthesis 903 Ribosomes have three tRNAbinding sites that bridge the 30s and 50s subunits 905 The promoter sites 865 RNA polymerases must unwind the start signal is usually AUG preceded by several bases that template pair with 16S rRNA 905 Bacterial protein synthesis is double helix for transcription to take place 865 Elongation initiated by takes place at transcription bubbles that move along the formylmethionyl transfer RNA 906 Formylmethionyl-tRNAf DNA template 866 Sequences within the newly transcribed is placed in the P site of the ribosome in the formation of RNA signal termination 866 Some messenger RNAs the 70S initiation complex 907 directly sense metabolite Elongation factors deliver aminoacyl-tRNA to the ribosome concentrations 867 The rho protein helps to terminate the 907 Peptidyl transferase catalyzes peptide-bond synthesis transcription of some genes 868 Some antibiotics inhibit 908 The formation of a peptide bond is followed by the transcription 869 Precursors of transfer and ribosomal RNA GTP driven translocation of tRNAs and mRNA 909 Protein are cleaved and chemically modified after transcription in synthesis is terminated by release factors that read stop prokaryotes 870 codons 910 29.2 Transcription in Eukaryotes Is Highly Regulated 871 Three 30.4 Eukaryotic Protein Synthesis Differs from Bacterial Protein types of RNA polymerase synthesize RNA in eukaryotic Synthesis Primarily in Translation Initiation 911 Mutations in cells 872 Three common elements can be found in the RNA polymerase II promoter region 874 The TFIID protein initiation factor 2 cause a curious pathological condition 913 30.5 A Variety of Antibiotics and complex initiates the assembly of the active transcription Toxins Can Inhibit Protein Synthesis 913 complex 874 Multiple transcription factors interact with Some antibiotics inhibit protein synthesis 914 Diphtheria eukaryotic promoters 875 Enhancer sequences can stimulate transcription at toxin blocks protein synthesis in start sites thousands of bases away 876 eukaryotes by inhibiting translocation 914 Ricin fatally 29.3 The Transcription Products of Eukaryotic Polymerases Are modifies 28S ribosomal RNA 915 30.6 Ribosomes Bound to the 859CHAPTER 29 RNA Synthesis and Processing 859 Processed 876 RNA polymerase I produces three ribosomal Endoplasmic Reticulum Manufacture Secretory and Membrane Proteins 915 Protein synthesis begins on ribosomes that are RNAs 877 RNA polymerase III produces transfer RNA 877 The product of RNA polymerase II, the pre-mRNA transcript, free in the cytoplasm 916 Signal sequences mark proteins for translocation acquires a 59 cap and a 39 poly(A) tail 878 Small regulatory across the endoplasmic reticulum membrane 916 RNAs are cleaved from larger precursors 879 RNA editing changes the proteins encoded by mRNA 879 Sequences at the ends of introns specify splice sites in mRNA precursors 880 Splicing consists of two sequential transesterification Transport vesicles carry cargo proteins to their final reactions 881 Small nuclear RNAs in spliceosomes catalyze the splicing of mRNA precursors 882 Transcription and processing of mRNA are coupled 883 Mutations that affect pre-mRNA splicing cause disease 884 Most human predestination 918 mRNAS can be spliced in alternative ways to yield different proteins 885 29.4 The Discovery of Catalytic RNA was Revealing in Regard CHAPTER 31 The Control of Gene Expression CHAPTER 31 The Control of Gene Expression in Prokaryotes 925 Controlled at Posttranscriptional Levels 954 Genes associated in Prokaryotes 925 31.1 Many DNA-Binding Proteins Recognize Specific DNA Sequences 926 The helix-turn-helix motif is common to many prokaryotic DNA-binding proteins 927 31.2 Prokaryotic with iron metabolism are translationally regulated in animals 954 Small RNAs regulate the expression of many eukaryotic genes 956 Part IV RESPONDING TO ENVIRONMENTAL DNA-Binding Proteins Bind Specifically to Regulatory Sites in Operons 927 An operon consists of regulatory elements and CHANGES protein-encoding genes 928 The lac repressor protein in the absence of lactose CHAPTER 33 Sensory Systems 961 CHAPTER 33 Sensory binds to the operator and blocks transcription 929 Ligand Systems 961 binding can induce structural changes in regulatory proteins 930 The operon is a common regulatory unit in prokaryotes 930 Transcription can be stimulated by 33.1 A Wide Variety of Organic Compounds Are Detected by proteins that contact RNA polymerase 931 Olfaction 962 Olfaction is mediated by an enormous family of seven-transmembrane-helix receptors 962 Odorants are 31.3 Regulatory Circuits Can Result in Switching Between Patterns of Gene Expression 932 The l repressor regulates its decoded by a combinatorial mechanism 964 33.2 Taste Is a own expression 932 A circuit based on the l repressor and Cro Combination of Senses That Function by Different Mechanisms forms 966 Sequencing of the human genome led to the discovery of a genetic switch 933 Many prokaryotic cells release a large family of 7TM bitter receptors 967 A heterodimeric chemical signals that regulate gene expression in other 7TM receptor responds to sweet cells 933 Biofilms are complex communities of prokaryotes compounds 968 Umami, the taste of glutamate and 934 aspartate, is mediated by a heterodimeric receptor related 31.4 Gene Expression Can Be Controlled at Posttranscriptional to the sweet receptor 969 Salty tastes are detected primarily by the passage of sodium ions through channels Levels 935 Attenuation is a prokaryotic mechanism for regulating transcription through the modulation of nascent 969 Sour tastes arise from the effects of hydrogen ions (acids) on channels 969 RNA secondary structure 935 CHAPTER 32 The Control of 33.3 Photoreceptor Molecules in the Eye Detect Visible Light 970 Rhodopsin, a specialized 7TM receptor, absorbs Gene Expression CHAPTER 32 The Control of Gene Expression in Eukaryotes 941 in Eukaryotes 941 visible light 970 Light absorption induces a specific isomerization of bound 11-cis-retinal 971 Light-induced lowering of the calcium level coordinates recovery 972 Color vision is mediated by three cone receptors that are homologs of rhodopsin 973 Rearrangements in the genes for the green and red pigments lead to “color blindness” 974 33.4 Hearing Depends on the Speedy Detection of Mechanical 32.1 Eukaryotic DNA Is Organized into Chromatin 943 Stimuli 975 Hair cells use a connected bundle of stereocilia to Nucleosomes are complexes of DNA and histones 943 DNA detect tiny motions 975 Contents xxxi wraps around histone octamers to form nucleosomes 943 32.2 Transcription Factors Bind DNA and Regulate Transcription Initiation 945 A range of DNA-binding structures are employed by eukaryotic DNA-binding proteins 945 Activation domains interact with other proteins 946 Multiple transcription factors interact with eukaryotic regulatory regions 946 Enhancers can stimulate transcription in specific cell types 946 Induced pluripotent stem cells can be generated by introducing four transcription factors into differentiated cells 947 32.3 The Control of Gene Expression Can Require Chromatin Remodeling 948 Mechanosensory channels have been identified in Drosophila and vertebrates 976 33.5 Touch Includes the Sensing of Pressure, Temperature, and Other Factors 977 Studies of capsaicin reveal a receptor for sensing high temperatures and other painful stimuli 977 CHAPTER 34 The Immune System 981CHAPTER 34 The Immune System 981 The methylation of DNA can alter patterns of gene expression 949 Steroids and related hydrophobic molecules pass through membranes and bind to DNAbinding receptors 949 Nuclear hormone receptors regulate transcription by recruiting coactivators to the transcription complex 950 Steroid-hormone receptors are targets for drugs 951 Chromatin structure is modulated through covalent modifications of histone tails 952 Histone deacetylases contribute to transcriptional repression 953 32.4 Eukaryotic Gene Expression Can Be Innate immunity is an evolutionarily ancient defense system 982 The adaptive immune system responds by using the principles of evolution 984 34.1 Antibodies Possess Distinct Antigen-Binding and Effector Units 985 34.2 Antibodies Bind Specific Molecules Through Hypervariable Loops 988 The immunoglobulin fold consists of a beta-sandwich framework with hypervariable loops 988 Xray analyses have revealed how antibodies bind antigens 989 Large antigens bind antibodies with numerous CHAPTER 36 Drug Development 1033 interactions 990 34.3 Diversity Is Generated by Gene Rearrangements 991 J (joining) genes and D (diversity) genes increase antibody diversity 991 More than 108 antibodies can be formed by combinatorial association and somatic mutation 992 The oligomerization of antibodies expressed on the surfaces of immature B cells triggers antibody secretion 993 Different classes of antibodies are formed by the hopping of VH genes 994 34.4 Major-Histocompatibility- Complex Proteins Present Peptide Antigens on Cell Surfaces for Recognition by T-Cell Receptors 995 Peptides presented by MHC proteins occupy a deep groove flanked by alpha helices 996 T-cell receptors are antibody-like proteins containing variable and constant regions 998 CD8 on cytotoxic T cells acts in concert with Tcell receptors 998 Helper T cells stimulate cells that display foreign peptides bound to class II MHC proteins 1000 Helper T cells rely on the T-cell receptor and CD4 to recognize foreign peptides on antigen-presenting cells 1000 MHC proteins are highly diverse 1002 Human immunodeficiency viruses subvert the immune system by destroying helper T cells 1003 34.5 The Immune System Contributes to the Prevention and the Development of Human Diseases 1004 T cells are subjected to positive and negative selection in the thymus 1004 Autoimmune diseases result from the generation of immune responses against self-antigens 1005 xxxii Contents 36.1 The Development of Drugs Presents Huge Challenges 1034 Drug candidates must be potent and selective modulators of their targets 1035 Drugs must have suitable properties to reach their targets 1036 Toxicity can limit drug effectiveness 1040 36.2 Drug Candidates Can Be Discovered by Serendipity, Screening, or Design 1041 Serendipitous observations can drive drug development 1041 Natural products are a valuable source of drugs and drug leads 1043 Screening libraries of synthetic compounds expands the opportunity for identification of drug leads 1044 Drugs can be designed on the basis of three-dimensional structural information about their targets 1046 36.3 Analyses of Genomes Hold Great Promise for Drug Discovery 1048 Potential targets can be identified in the human proteome 1048 Animal models can be developed to test the validity of potential drug targets 1049 Potential targets can be identified in the genomes of pathogens 1050 Genetic differences influence individual responses to drugs 1050 36.4 The Clinical Development of Drugs Proceeds Through Several Phases 1051 Clinical trials are time consuming and expensive 1052 The evolution of drug resistance can limit the utility of drugs for infectious agents and cancer 1053 Answers to Problems A1 Selected Readings B1 Index C1 The immune system plays a role in cancer prevention 1005 Vaccines are a powerful means to prevent and eradicate disease 1006 CHAPTER 35 Molecular Motors 1011 CHAPTER 35 Molecular Motors 1011 35.1 Most Molecular-Motor Proteins Are Members of the PLoop NTPase Superfamily 1012 Molecular motors are generally oligomeric proteins with an ATPase core and an extended structure 1012 ATP binding and hydrolysis induce changes in the conformation and binding affinity of motor proteins 1014 35.2 Myosins Move Along Actin Filaments 1016 Actin is a polar, self-assembling, dynamic polymer 1016 Myosin head domains bind to actin filaments 1018 Motions of single motor proteins can be directly observed 1018 Phosphate release triggers the myosin power stroke 1019 Muscle is a complex of myosin and actin 1019 The length of the lever arm determines motor velocity 1022 35.3 Kinesin and Dynein Move Along Microtubules 1022 Microtubules are hollow cylindrical polymers 1022 Kinesin motion is highly processive 1024 35.4 A Rotary Motor Drives Bacterial Motion 1026 Bacteria swim by rotating their flagella 1026 Proton flow drives bacterial flagellar rotation 1026 Bacterial chemotaxis depends on reversal of the direction of flagellar rotation 1028 CHAPTER 36 Drug Biochemistry: An Evolving Science 1 CHAPTER C H O H2 C HN OC CH 2 + H+ C O — C H O H2 C HN OC CH C 2 O H Chemistry in action.Human activities require energy. The interconversionof different forms of energy requires large biochemical machines comprising many thousands of atoms such as the complex shown above. Yet, the functions of these elaborate assemblies depend on simple chemical processes such as the protonationand deprotonationof the carboxylic acid groups shown on the right. The photograph is of Nobel Prize winners Peter Agre , M.D., and Carol Greider , Ph.D., who used, respectively, biochemical techniques to reveal key mechanisms of how water is transported into and out of cells, and how chromosomes are replicated faithfully. [Keith Weller for Johns Hopkins Medicine.] B iochemistry is the study of the chemistry of life processes. Since the the great unity of all living things at the biochemical level. 1.1 Biochemical Unity Underlies Biological discovery that biological molecules such as urea Diversity could be synthesized from nonliving components in 1828, scientists have explored the chemistry of The biological world is magnificently diverse. The animal kingdom is rich with species ranging from life with great intensity. Through these nearly microscopic insects to elephants and investigations, many of the most fundamental whales. The plant kingdom includes species as mysteries of how living things function at a small and relatively biochemical level have now been solved. However, much remains to be investigated. As is often the case, each discovery raises at least as O U T L I N E many new questions as it answers. Furthermore, 1.1 Biochemical Unity Underlies Biological Diversity we are now in an age of unprecedented opportunity for the application of our tremendous 1.2 DNA Illustrates the Interplay BetweenForm and Function knowledge of biochemistry to problems in 1.3 Concepts from Chemistry Explain the Properties of medicine, dentistry, agriculture, forensics, Biological Molecules anthropology, envi ronmental sciences, alternative energy, and many other fields. We begin our 1.4 The Genomic Revolution Is Transforming Biochemistry, journey into biochemistry with one of the most Medicine, and Other Fields startling discoveries of the past century: namely, simple as algae and as large and from cells suggested that these CHAPTER 1 Biochemistry: An Evolving Science complex as giant sequoias. This diverse organisms might have diversity extends further when we more in common than is apparent descend into the microscopic from their outward appearance. world. Organisms such as With the development of protozoa, yeast, and bacteria are biochemistry, this suggestion has present with great diversity in been tremendously supported and water, in soil, and on or within expanded. At the biochemical level, larger organisms. Some organisms all organisms have many common can survive and even thrive in features (Figure 1.1). seemingly hostile environments As mentioned earlier, biochemistry such as hot springs and glaciers. is the study of the chemistry of life The development of the processes. These processes entail microscope revealed a key unifying the interplay of two different feature that underlies this diversity. classes of molecules: large CH2OH Large organisms are built up of molecules such as proteins and O cells, resembling, to some extent, nucleic acids, referred to as single-celled microscopic biological macromolecules, and OH CH2OH organisms. The construction of ani low-molecular-weight molecules mals, plants, and microorganisms such as glu HO C H referred to as metabolites, things. For example, built from the same set of HO that are chemically deoxyribonucleic acid 20 building blocks in all OH transformed in biological (DNA) stores genetic organisms. Furthermore, OH processes. information in all cellular proteins that play similar Glucose Members of both these organisms. Proteins, the roles in different classes of molecules are macromol ecules that are organisms often have very CH2OH key participants in most similar three dimensional Glycerol common, with minor cose and glycerol, biological processes, are structures (Figure 1.1). variations, to all living 2 Sulfolobus archaea Arabidopsis thaliana Homo sapiens 1 FIGURE 1.1 Biological diversity and similarity. The shape of a key molecule in gene regulation (the TATA-box-binding protein) is similar in three very different organisms that are separated from one another by billions of years of evolution. [(Left) Eye of Science/Science Source; (middle) Holt Studios/Photo Researchers; (right) Time Life Pictures/Getty Images.] lc r a u m u n s m a i de s u c n ip o o m a r h ti s c a iM E ll w e i e s 3 m o r iD s s o ht n gr o f H s c r C Oxygen atmosphere forming o s r i g 1.1 Unity and Diversity n c a i n a e gr b o s n M 4.5 4.0 3.5 3.0 2.5 2.0 1.5 1.0 0.5 0.0 Billions of years FIGURE 1.2 A possible time line for biochemical evolution. Selected key events are indicated. Note that life on Earth began approximately 3.5 billion years ago, whereas human beings emerged quite recently. Key metabolic processes also are common to many organisms. For example, the set of chemical transformations that converts glucose and oxy gen into carbon dioxide and water is essentially identical in simple bacteria such as Escherichia coli (E. coli) and human beings. Even processes that appear to be quite distinct often have common features at the biochemical level. Remarkably, the biochemical processes by which plants capture light energy and convert it into more-useful forms are strikingly similar to steps used in animals to capture energy released from the breakdown of glucose. These observations overwhelmingly suggest that all living things on Earth have a common ancestor and that modern organisms have evolved from this ancestor into their present forms. Geological and biochemical find ings support a time line for this evolutionary path (Figure 1.2). On the basis of their biochemical characteristics, the diverse organisms of the modern world can be divided into three fundamental groups called domains: Eukarya (eukaryotes), Bacteria, and Archaea . Domain Eukarya comprises all multicel lular organisms, including human beings as well as many microscopic unicel lular organisms such as yeast. The defining characteristic of eukaryotes is the presence of a well-defined nucleus within each cell. Unicellular organisms such as bacteria, which lack a nucleus, are referred to as prokaryotes . The pro karyotes were reclassified as two separate domains in response to Carl Woese’s discovery in 1977 that certain bacteria-like organisms are biochemi cally quite distinct from other previously characterized bacterial species. These organisms, now recognized as having diverged from bacteria early in evolution, are the archaea . Evolutionary paths from a common ancestor to modern organisms can BACTERIA EUKARYA ARCHAEA be deduced on the basis of biochemical information. One s s s such path is shown in Figure 1.3. e u m u c c u b c y i Much of this book will explore the chemical reactions a a r o o l l i m e c l t g h o o e c r c s o and the associated biological macromolecules and metab i n n a a e r u a o o l b h e a l h i c o olites that are found in biological processes common to all h h t m l m c a l c c c e o a a r a a s e E S B H S Z M A H organisms. The unity of life at the biochemical level makes this approach possible. At the same time, different organisms have specific needs, depending on the particu lar biological niche in which they evolved and live. By comparing and contrasting details of particular biochemi cal pathways in different organisms, we can learn how biological challenges are solved at the biochemical level. In most cases, these challenges are addressed by the adap tation of existing macromolecules to new roles rather than by the evolution of entirely new ones. Biochemistry has been greatly enriched by our for many of the advances in biochemistry and many ability to examine the three-dimensional structures of other fields, extending to the present. biological macromolecules in great detail. Some of The structure of DNA powerfully illustrates a basic these structures principle common to all biological macromolecules: the intimate relation between structure and function. FIGURE 1.3 The tree of life. A possible evolutionary path from a common The remarkable properties of this chemical ancestor approximately 3.5 billion years ago at the bottom of the tree to substance allow it to function as a very efficient and organisms found in the modern world at the top. 4 robust vehicle for storing information. We start with an examination of the covalent structure of DNA and CHAPTER 1 Biochemistry: An Evolving Science its exten are simple and elegant, whereas others are incredibly complicated. In any case, these structures sion into three dimensions. provide an essential framework for understanding function. We begin our exploration of the interplay DNA is constructed from four building blocks between structure and function with the genetic DNA is a linear polymer made up of four different material, DNA. types of monomers. It has a fixed backbone from which protrude variable substituents , referred to as bases (Figure 1.4). The backbone is built of 1.2 DNA Illustrates the Interplay Between repeating sugar–phosphate units. The sugars are Form and Function molecules of deoxyribose from which DNA receives its name. Each sugar is connected to two phosphate A fundamental biochemical feature common to all groups through different linkages. Moreover, each cellular organisms is the use of DNA for the storage sugar is oriented in the same way, and so each DNA of genetic information. The discovery that DNA plays strand has directionality, with one end distinguishable this central role was first made in studies of bacteria from the other. Joined to each deoxyribose is one of in the 1940s. This discovery was followed by a four possible bases: adenine (A), cyto sine (C), compelling proposal for the three-dimensional guanine (G), and thymine (T). structure of DNA in 1953, an event that set the stage NH2 N NH2 H N N H N H NN O N N H N O O NH H CH3 N H NH2 O NH Adenine (A) Cytosine (C) Guanine (G) Thymine (T) These bases are connected to the sugar components in the DNA backbone through the bonds shown in black in Figure 1.4. All four bases are planar but differ significantly in other respects. Thus, each monomer of DNA consists of a sugar–phosphate unit and one of four bases attached to the sugar. These bases can be arranged in any order along a strand of DNA. base1 base2 base3 OO O FIGURE 1.4 Covalent structure of O DNA. Each unit of the polymeric structure O OO OO is composed of a sugar (deoxyribose), a phosphate, and a variable base that O P P P OO protrudes from the sugar–phosphate backbone. O O –– – OO Sugar Phosphate Two single strands of DNA combine to form a double helix Most DNA molecules consist of not one but two strands (Figure 1.5). In 1953, James Watson and Francis Crick deduced the arrangement of these strands and proposed a three-dimensional structure for DNA molecules. This structure is a carbon or carbon–nitrogen bonds that define the double helix composed of two inter struc tures of the bases themselves. Such weak twined strands arranged such that the sugar– bonds are crucial to biochemical systems; they phosphate backbone lies on the outside and the are weak enough to be reversibly broken in bases on the inside. The key to this structure is biochemical processes, yet they are strong that the bases form specific base pairs ( bp ) held enough, particularly when many form simultane together by hydrogen bonds (Section 1.3): ously, to help stabilize specific structures such as adenine pairs with thymine (A–T) and guanine the double helix. pairs with cytosine (G–C), as shown in Figure 1.6. Hydrogen bonds are much weaker than covalent FIGURE 1.5 The double helix. The double-helical structure of DNA bonds such as the carbon– proposed by Watson and Crick. The sugar–phosphate backbones of 5 the two chains are shown in red and blue, and the bases are shown in green, purple, orange, and yellow. The two strands are antiparallel, running in opposite directions with respect to the axis of the double helix, as indicated by the arrows. 1.2 DNA: Form and Function H N H O N N N H N NHN H CH3 N O N N H N O N N N O N H H Guanine (G) Cytosine (C) Adenine (A) Thymine (T) shape (Figure 1.6) and thus fit equally well into the center of the double-helical (A – T), and guanine with cytosine (G – C). The dashed green structure of any sequence. Without any lines represent hydrogen bonds. constraints, the sequence of bases along a DNA strand can act as an efficient means of storing information. Indeed, the DNA structure explains heredity and the storage of sequence of bases along DNA strands is information how genetic information is stored. The structure proposed by Watson and Crick has two properties of central importance to the role of DNA as the hereditary material. First, the structure is compatible with any sequence of bases. While the bases are distinct in structure, the base pairs have essentially the same FIGURE 1.6 Watson–Crick base pairs. Adenine pairs with thymine G C A sequences and protein activities base-pairing, one strand C C The DNA of the molecules within cells. the Newly sequence ribonucleic that carry out Second, sequence of synthesized determines acid (RNA) most of the because of bases along strands C A G the sequence along G wrote: “It has not specific pairing the other strand. GG escaped our completely Crick so coyly As Watson and C determines the notice that the T T C T A T A A G we have postulated immediately suggests a possible copying mechanism for T the genetic material.” Thus, if the DNA double helix is separated into two G C 1.3 Concepts from Chemistry Explain the Properties of Biological Molecules single strands, each strand can act as a template for the generation of its partner strand through specific base-pair formation (Figure 1.7). The three We have seen how a chemical insight into the dimensional structure of DNA beautifully illustrates hydrogen-bonding capabili ties of the bases of DNA the close connection between molecular form and led to a deep understanding of a fundamental function. biological process. To lay the groundwork for the rest of the book, we begin our study of FIGURE 1.7 DNA replication. If a DNA molecule is separated into two strands, each strand can act as the template for the generation of its partner biochemistry by examining selected concepts from strand. chemistry and showing how these concepts apply 6 to biological systems. The concepts include the CHAPTER 1 Biochemistry: types of chemical bonds; the structure of water, the An Evolving Science solvent in which most biochemical processes take place; the First and Second Laws of Thermodynamics; and the principles of acid–base chemistry. The formation of the DNA double helix as a key example We will use these concepts to examine an archetypical biochemical process— namely, the formation of a DNA double helix from its two component strands. The process is but one of many examples that could have been chosen to illustrate these topics. Keep in mind that, although the specific discussion is about DNA and doublehelix formation, the concepts consid ered are quite general and will apply to many other classes of molecules and processes that will be discussed in the remainder of the book. In the course of these discussions, we will touch on the properties of water and the concepts of pK a and buffers that are of great importance to many aspects of biochemistry. The double helix can form from its component strands G CG C GATTAAT CTAATTA GATTAAT CGATTAAT and The discovery that DNA from natural sources exists in a double-helical form with Watson–Crick base pairs suggested, but did not prove, that such double helices would form spontaneously outside biological systems. Suppose that two short strands of DNA were chemically synthesized to have complementary sequences so that they could, in principle, form a double helix with Watson–Crick base pairs. Two such sequences are molecules C in solution can be sequences are mixed, a What forces cause the examined by a variety of double helix with T two strands of DNA to techniques. In isolation, Watson–Crick base pairs bind to each other? T does form (Figure 1.8). each sequence exists This reaction pro almost exclusively as a T single-stranded molecule. ceeds nearly to A completion. A A However, when the two several factors: the types of interactions and bonds FIGURE 1.8 Formation of a double helix. When two DNA strands with in biochemical systems and the ener getic favorability appropriate, complementary sequences are mixed, they spontaneously assemble of the reaction. We must also consider the influence to form a double helix. To analyze this binding reaction, we must consider of the solution conditions—in particular, the consequences of acid– base reactions. Covalent and noncovalentbonds are important for the structure and stability of biological molecules Atoms interact with one another through chemical bonds. These bonds include the covalent bonds that define the structure of molecules as well as a variety of noncovalent bonds that are of great importance to biochemistry. Covalent bonds. The strongest bonds are covalent bonds, such as the bonds that hold the atoms together within the individual bases shown on page 4. A covalent bond is formed by the sharing of a pair of electrons between adjacent atoms. A typical carbon–carbon (C}C) covalent bond has a bond length of 1.54 Å and bond energy of 355 kJ covalent bonding can be written. For example, mol 1 (85 kcal mol 1 ). Because covalent bonds are adenine can be written in two nearly equivalent ways so strong, considerable energy must be expended to called resonance structures. 7 break them. More than one electron pair can be 1.3 Chemical Concepts shared between two atoms to form a multiple covalent bond. For example, three of the bases in Figure 1.6 include carbon–oxygen (C“O) double Distance and energy units bonds. These bonds are even stronger than C}C Interatomic distances and bond lengths are usually measured in angstrom (Å) units: single bonds, with energies near 730 kJ mol 1 1 Å 5 10210 m 5 1028 cm 5 0.1 nm (175 kcal mol 1 ) and are somewhat shorter. Several energy units are in common use. One joule (J) is the amount of energy For some molecules, more than one pattern of required to move 1 meter against a force of 1 newton. required to raise the NH2 temperature of 1 gram of water 1 5 degree A kilojoule (kJ) is Celsius. A kilocalorie One joule is equal to 1000 joules. One (kcal) is 1000 0.239 cal. calorie is calories. the amount of energy N NH2 5 H NN H NN N 4N N H 4 H stability than does a molecule without multiple resonance structures. These adenine structures depict alternative arrangements of single and double bonds that Noncovalent bonds. Noncovalent bonds are are possible within the same structural weaker than covalent bonds but are crucial for framework. Resonance structures are shown biochemical processes such as the formation of a connected by a double-headed arrow. Adenine’s double helix. Four fundamental noncovalent true structure is a composite of its two resonance bond types are ionic interactions, hydrogen structures. The composite structure is bonds, van der Waals interactions, and manifested in the bond lengths such as that for hydrophobic interactions . They differ in the bond joining carbon atoms C-4 and C-5. The geometry, strength, and specificity. Furthermore, observed bond length of 1.40 Å is between that these bonds are affected in vastly different ways expected for a C}C single bond (1.54 Å) and a by the presence of water. Let us consider the C“C double bond (1.34 Å). A molecule that can characteristics of each type: be written as several resonance structures of approximately equal energies has greater 1. Ionic Interactions . A charged group on one q1 q2 r molecule can attract an oppo sitely charged group on the same or another molecule. The energy of an ionic interaction (sometimes called 8 CHAPTER 1 Biochemistry: An Evolving Science an electrostatic interaction) is given by the 2. Hydrogen Bonds . These interactions are Coulomb energy: largely ionic interactions, with partial charges on nearby atoms attracting one another. Hydrogen E 5 kq1q2/Dr bonds are responsible for specific base-pair where E is the energy, q1 and q2 are the charges formation in the DNA double helix. The hydrogen on the two atoms (in units of the electronic atom in a hydrogen bond is partially shared by charge), r is the distance between the two atoms two electronega (in ang stroms), D is the dielectric constant tive atoms such as nitrogen or oxygen. The (which decreases the strength of the Coulomb hydrogen-bond donor is the group depending on the intervening solvent or medium), and k is a pro portionality constant (k 5 1389, for energies in units of kilojoules per mole, or 332 for energies in kilocalories per mole). By convention, an attractive interaction has a negative energy. The ionic interaction between two ions bearing single opposite charges sepa rated by 3 Å in water (which has a dielectric constant of 80) has an energy of 2 5.8 kJ mol 1 ( 2 1.4 kcal mol 1 ). Note how important the dielectric constant of the medium is. For the same ions separated by 3 Å in a nonpolar solvent such as hexane (which has a dielectric constant of 2), the energy of this interaction is 2 232 kJ mol 1 ( 2 55 kcal mol 1 ). Hydrogen bond donor that includes both the atom to atom itself, whereas the Hydrogen which the hydrogen atom is more hydrogen-bond acceptor is bond acceptor tightly linked and the hydrogen NHN develops a partial positive charge ( d ). Thus, the + − − hydrogen atom with a partial positive charge can interact with an atom having a partial negative NHO charge ( d ) through an ionic interaction. Hydrogen bonds are much weaker than covalent OHN bonds. They have ener gies ranging from 4 to 20 kJ OHO mol 1 (from 1 to 5 kcal mol 1 ). Hydrogen bonds are FIGURE 1.9 Hydrogen bonds. Hydrogen bonds are depicted by dashed green also somewhat longer than covalent bonds; their lines. The positions of the partial charges (d and d ) are shown. bond lengths (measured from the hydrogen atom) the atom less tightly linked to the hydrogen atom range from 1.5 Å to 2.6 Å; hence, a distance ranging (Figure 1.9). The electro negative atom to which the from 2.4 Å to 3.5 Å separates the two nonhydrogen hydrogen atom is covalently bonded pulls elec tron atoms density away from the hydrogen atom, which thus approximately straight, such that in a hydrogen bond. Hydrogen bond donor The strongest hydrogen bonds the hydrogen-bond donor, the hydrogen atom, and the have a tendency to be 0.9 Å 2.0 Å Hydrogen-bond acceptor t c a r t t A N H O 180° y gr e n E n o i s lu p e R n o i 0 van der Waals respect to one fluctuates with time. another. Hydrogen- At any instant, the bonding interactions charge distribution is are responsible for not perfectly many of the symmetric. This properties of water transient asymmetry that make it such a in the electronic special solvent, as charge about an will be described atom acts through contact distance Distance shortly. ionic interactions to induce a hydrogen-bond 3. van der Waals complementary acceptor lie along a Interactions . The asymmetry in the straight line. This basis of a van der tendency toward lin Waals interaction is electron distribution within its neighboring earity can be that the distribution atoms. The atom and important for of electronic charge its neighbors then orienting interacting around an atom attract one another. molecules with This attraction increases as two atoms come closer to each other, until they are separated distances shorter repulsive forces the two atoms by the van der Waals than the van der become dominant overlap. contact distance Waals contact because the outer (Figure 1.10). At distance, very strong electron clouds of FIGURE 1.10 Energy of a van der Waals interaction as two atoms approach each hydrogen bonds hold the structure together; similar other. The energy is most favorable at the van der Waals contact distance. interactions link molecules in Owing to electron–electron repulsion, the energy rises rapidly as the distance liquid water and account for many of the properties of between the atoms becomes shorter than the contact distance. water. In the liquid state, approximately one in four of the hydrogen bonds present in ice are broken. The polar nature of water is responsible for its high dielectric constant of 80. Molecules in aqueous solution interact with water mole cules through the formation of hydrogen bonds and through ionic interactions. These interactions make water a versatile solvent, able to readily dissolve many species, especially polar and charged Electric dipole compounds that can participate in these interactions. Energies associated with van der Waals interactions are quite small; typical interactions contribute from 2 The hydrophobic effect. A final fundamental to 4 kJ mol 1 (from 0.5 to 1 kcal mol 1 ) per atom interaction called the hydrophobic effect is a pair. When the surfaces of two large molecules come manifestation of the proper ties of water. Some together, however, a large number of atoms are in molecules (termed nonpolar molecules ) cannot van der Waals contact, and the net effect, summed participate in hydrogen bonding or ionic interac over many atom pairs, can be substantial. 9 We will cover the fourth noncovalent interaction, the 1.3 Chemical Concepts hydrophobic inter action, after we examine the characteristics of water; these characteristics are essential to understanding the hydrophobic interaction. Properties of water. Water is the solvent in which most biochemical reac tions take place, and its properties are essential to the formation of macro molecular structures and the progress of chemical reactions. Two properties of water are especially relevant: 1. Water is a polar molecule . The water molecule is bent, not linear, and so O – the distribution of charge is asymmetric. The oxygen nucleus draws elec HH+ trons away from the two hydrogen nuclei, which leaves the region around tions. The interactions of nonpolar molecules with each hydrogen atom with a net positive charge. The water molecules are not as favorable as are water molecule is thus an electrically polar structure. interactions between the water molecules themselves. The water molecules in contact with 2. Water is highly cohesive . Water molecules interact strongly with one another through hydrogen these nonpolar molecules form “cages” around bonds. These interactions are apparent in FIGURE 1.11 Structure of ice. Hydrogen bonds (shown as dashed green lines) are formed between water molecules to produce a highly ordered and open the structure of ice (Figure 1.11). Networks of structure. them, becoming more well ordered than water molecules free in solution. However, when two such nonpolar molecules come together, some of the water molecules are released, allowing them to interact freely with bulk water (Figure 1.12). The release of water from such cages is favorable for reasons to be considered shortly. The result is that nonpolar molecules show an increased tendency to associate with one another in water compared with other, less polar and less self-associating, solvents. This tendency is called the hydrophobic effect and the associated interactions are called hydrophobic interactions . The double helix is an expression of the rules of chemistry Let us now see how these four noncovalent interactions work together in driving the association of two strands of DNA to form a double helix. First, each phosphate group in a DNA strand carries a negative charge. These negatively charged groups interact unfavorably with one another over dis tances. Thus, unfavorable ionic interactions take place when two strands of Nonpolar molecule molecule Nonpolar molecule aggregation of nonpolar groups in water leads to the release of water molecules, initially interacting with the nonpolar surface, into bulk water. The release of water molecules into solution makes the aggregation of nonpolar groups favorable. Nonpolar molecule Nonpolar FIGURE 1.12 The hydrophobic effect. The DNA come together. These phosphate groups are far apart in the double helix with distances greater than 10 Å, but many such interactions take place (Figure 1.13). Thus, ionic interactions oppose the formation of the double helix. The strength of these repulsive ionic interactions is dimin ished by the high dielectric constant of water and the presence of ionic species such as Na or Mg 2 ions in solution. These positively charged species interact with the phosphate groups and partly neutralize their negative charges. Second, as already noted, hydrogen bonds are important in determining the formation of specific base pairs in the double helix. However, in single stranded DNA, the hydrogen-bond donors and acceptors are exposed to solution and can form hydrogen bonds with water molecules. FIGURE 1.13 Ionic (the phosphorus atom being O shown in purple) that bears H a negative charge. The H C interactions in DNA. Each O unit within the double helix includes a phosphate group H + O O H H H OH HN C H O H N unfavorable interactions of one phosphate with several others are shown by red lines. These repulsive interactions oppose the formation of a double helix. FIGURE 1.14 Base stacking. In the DNA double helix, adjacent base pairs are stacked nearly on top of one another, and so many atoms in each base pair are separated by their van der Waals contact distance. The central base pair is shown in dark blue and the two adjacent base pairs in light blue. Several van der Waals contacts are shown in red. van der Waals contacts When two single strands come together, these hydrogen bonds with water are broken and new hydrogen bonds between the bases are formed. Because the number of hydrogen bonds broken is the same as the number formed, these hydrogen bonds do not contribute substantially to driving the overall process of double-helix formation. However, favorability of base stacking. More-complete base they contribute greatly to the specificity of binding. stacking moves the nonpolar surfaces of the bases Suppose two bases that cannot form Watson–Crick out of water into contact with each other. base pairs are brought together. Hydrogen bonds The principles of double-helix formation between two with water must be bro strands of DNA apply to many other biochemical ken as the bases come into contact. Because the processes. Many weak interactions con tribute to the bases are not complemen tary in structure, not all of overall energetics of the process, some favorably and some unfavorably. Furthermore, surface these bonds can be simultaneously replaced by hydrogen bonds between the bases. Thus, the complementarity is a key feature: when formation of a double helix between complementary surfaces meet, hydrogen-bond noncomplementary sequences is disfavored. donors align with hydrogen bond acceptors and nonpolar surfaces come together to maximize van Third, within a double helix, the base pairs are parallel and stacked nearly on top of one another. der Waals interactions and minimize nonpolar surface area exposed to the aque ous environment. The typical separation between the planes of The properties of water play a major role in adjacent base pairs is 3.4 Å, and the distances determining the importance of these interactions. between the most closely approaching atoms are approximately 3.6 Å. This separation distance cor The laws of thermodynamics govern the behavior of responds nicely to the van der Waals contact distance (Figure 1.14). Bases tend to stack even in biochemical systems single-stranded DNA molecules. However, the base We can look at the formation of the double helix from stacking and associated van der Waals interactions a different perspec tive by examining the laws of are nearly optimal in a double-helical structure. thermodynamics. These laws are general Fourth, the hydrophobic effect also contributes to the 10 principles that apply to all physical (and biological) system plus that of its sur roundings always processes. They are of great importance because increases. For example, the release of water from they determine the conditions under which spe cific nonpolar surfaces responsible for the hydrophobic processes can or cannot take place. We will consider effect is favorable because water molecules free in these laws from a general perspective first and then solution are more disordered than they are when apply the principles that we have devel oped to the they are associated with nonpolar surfaces. At first formation of the double helix. glance, the Second Law appears to contradict much The laws of thermodynamics distinguish between a common experience, particularly about biological sys system and its surroundings. A system refers to the tems. Many biological processes, such as the matter within a defined region of space. The matter generation of a leaf from car bon dioxide gas and in the rest of the universe is called the surroundings. other nutrients, clearly increase the level of order and hence decrease entropy. Entropy may be The First Law of Thermodynamics states that the decreased locally in the formation of such ordered total energy of a system and its surroundings is constant . In other words, the energy content of the structures only if the entropy of other parts of the universe is increased by an equal or greater uni verse is constant; energy can be neither created nor amount. The local decrease in entropy is often accomplished by a release of heat, which increases destroyed. Energy can take different forms, however. Heat, for example, is one form of energy. the entropy of the surroundings. We can analyze this process in quantitative terms. Heat is a manifestation of the kinetic energy First, consider the system. The entropy ( S ) of the associated with the random motion of molecules. system may change in the course of a chemical Alternatively, energy can be present as potential energy —energy that will be released on the reaction by an amount DSsystem . If heat flows from occurrence of some process. Consider, for example, the system to its surroundings, then the heat a ball held at the top of a tower. The ball has content, often referred to as the enthalpy ( H ) , considerable potential energy because, when it is of the system will be reduced by an amount released, the ball will develop kinetic energy DHsystem . To apply the Second Law, we must associated with its motion as it falls. Within chemical determine the change in entropy of the surround systems, potential energy is related to the likelihood ings. If heat flows from the system to the that atoms can react with one another. For instance, surroundings, then the entropy of the surroundings a mixture of gasoline and oxy gen has a large will increase. The precise change in the entropy of potential energy because these molecules may react the surroundings depends on the temperature; the to form carbon dioxide and water and release change in entropy is greater when heat is added to energy as heat. The First Law requires that any relatively cold surroundings than when heat is added energy released in the formation of chemical bonds to surroundings at high temperatures that are must be used to break other bonds, released as heat already in a high degree of disorder. To be even or light, or stored in some other form. more specific, the change in the entropy of the sur Another important thermodynamic concept is that of roundings will be proportional to the amount of heat entropy, a measure of the degree of randomness or transferred from the system and inversely disorder in a system. The Second Law of proportional to the temperature ( T ) of the surround Thermodynamics states that the total entropy of a ings. In biological systems, T [in kelvins (K), absolute temperature] is if and only if 11 ¢Ssystem . ¢HsystemyT (6) Rearranging gives 1.3 Chemical Concepts 12 TDSsystem . DH or, in other words, entropy will increase if and only if CHAPTER 1 Biochemistry: An Evolving Science usually assumed to be constant. Thus, a change in the entropy of the surroundings is given by ¢G 5 ¢Hsystem 2 T¢Ssystem , 0 (7) ¢Ssurroundings 5 2¢HsystemyT (1) The total entropy Thus, the free-energy change must be negative for a process to take place spontaneously. There is change is given by the expression negative free-energy change when and only when the overall entropy of the universe is increased . Again, the free energy represents a single term that takes into account both the entropy of the system and the entropy of the surroundings. Heat is released in the formation of the double helix ¢Stotal 5 ¢Ssystem 1 ¢Ssurroundings (2) Let us see how the principles of thermodynamics apply to the formation of the double helix (Figure Substituting equation 1 into equation 2 yields 1.15). Suppose solutions containing each of the two ¢Stotal 5 ¢Ssystem 2 ¢HsystemyT (3) Multiplying by single strands are mixed. Before the double helix forms, each of the single strands is free to translate 2T gives and rotate in solution, whereas each matched pair of strands in the double helix must move together. 2T¢Stotal 5 ¢Hsystem 2 T¢Ssystem (4) Furthermore, the free single strands exist in more The function 2TDS has units of energy and is referred to as free energy or Gibbs free energy, after conformations than possible when bound together in Josiah Willard Gibbs, who developed this function in a double helix. Thus, the formation of a double helix from two single strands appears to result in an 1878: increase in order for the system, that is, a decrease in the entropy of the system. ¢G 5 ¢Hsystem 2 T¢Ssystem (5) On the basis of this analysis, we expect that the The free-energy change, DG , will be used double helix cannot form without violating the throughout this book to describe the energetics of Second Law of Thermodynamics unless heat is biochemical reactions. The Gibbs free energy is released to increase the entropy of the essentially an accounting tool that keeps track of surroundings. Experimentally, we can measure the both the entropy of the system (directly) and the heat released by allowing the solutions containing entropy of the surroundings (in the form of heat the two single strands to come together within a released from the system). water bath, which here corresponds to the surround Recall that the Second Law of Thermodynamics ings. We then determine how much heat must be states that, for a process to take place, the entropy absorbed by the water bath or released from it to of the universe must increase. Examination of maintain it at a constant temperature. This equation 3 shows that the total entropy will increase experiment reveals that approximatelresult quite large, a substantial y 250 kJ mol reveals that 2 250 kJ mol amount of 1 (60 kcal the change 1 , consistent heat is mol 1 ). This in enthalpy with our released— experimentalfor the namely, process is C T A A C T A A T T A GCT AAT AATT A C T A GATT TTA AAT A CG C GATT GCT GAT T AA T AATT C A G G CG T T expectation that significant heat would have to released to the surroundings for the process be not to violate the roundings to Second Law. We ensure that the see in quanti entropy of the tative terms how uni verse order within a increases. We Acid–base reactions system can be will encounter increased by are central in many this general releasing theme again and sufficient heat to again throughout the sur this book. AG T A A T C G A T A TA T A A T T GC A T G C TAA A T TA T C A T T C G C GATT AAT T Mixing TTAA A CG GCT AATT C GATT G C T A T T A A A T T A A T C G A A T GCT C G A T T A A T G C T A TTA C G T TTA TAA biochemical processes Throughou t our considerati on of the formation AATT AAT TAA C G A T T TTAA A A A of the double helix, we have dealt only with the C A AATT GATT noncovalent bonds that are formed or broken in this bonds. A the formation particularly and cleavage important of covalent class of T T A A GCT process. Many biochemical processes entail A A GATT CG C G A T T G C T A CG A T AAT AAT C GAT T AA T C reactions reactions . promi nent in biochemistry is acid – base T C T A A T T A T A A T A A T T T GC AAT A A T T A A G C T C A T GCT A G are added to Throughout the which the In acid and basemolecules or book, we will addition or reactions, removed from encounter many removal of hydrogen ions them. processes in atoms is processes crucial, such by which C as the meta carbohydrat TA bolic es are hydrogen release energy for other uses. Thus, degraded to TT understanding of a thorough the basic princi reactions is essential. ples of these written as H , cor A hydrogen ion, often responds to a proton. In fact, hydrogen ions C G G A T T A TTA T TA A T A A T T A A A T GC A T T ATT A Reacting TTA TAA T TAA GA G C T C G A GC A C G AC C TTAA A TT A G CG G A A T T TT T A TTAA TTAA A A TAA TAG C T A C GCT TAA G T AATT TC A GATT GCT C AATT T AAT A GAT A A GCT C T exist in solution bound to water molecules, thus T A AAT GATT T T C A T G AATT forming what are known as hydronium ions, H 3 O . A A A G T T G C T C T C A AAT A A G For simplicity, we will continue to write H , A T T A T A but we should keep in mind that H is short hand for the actual species present. The concentration of hydrogen ions in solu tion is expressed as the pH. Specifically, the pH of a solution is defined as FIGURE 1.15 Double-helix formation and entropy. When solutions containing DNA strands with complementary sequences are mixed, the strands react to form double helices. This process results in a loss of entropy from the system, indicating that heat must be released to the surroundings to prevent a violation of the Second Law of Thermodynamics. pH 5 2log[H1] where [H ] is in units of molarity. Thus, pH 7.0 refers to a solution for which 2 log[H ] 5 7.0, and so log[H ] 5 2 7.0 and [H ] 5 10 log[H ] 5 10 7.0 5 1.0 3 10 7 M. The pH also indirectly expresses the concentration of hydroxide ions, [OH ], in solution. To see how, we must realize that water molecules can dissociate to form H and OH ions in an equilibrium process. H2O Δ H1 1 OH2 The equilibrium constant ( K ) for the dissociation of water is defined as K 5 [H1][OH2]/[H2O] 13 14 The concentration of water, [H 2 O], in pure water is 55.5 M, and this concentration is constant under most conditions. Thus, we can define a new constant, KW : CHAPTER 1 Biochemistry: An Evolving Science KW 5 K[H2O] 5 [H1][OH2] K[H2O] 5 1.8 3 10216 3 55.5 5 1.0 3 10214 Because KW 5 [H ][OH ] 5 1.0 3 10 14 , we can calculate [OH2] 5 10214/[H1] and [H1] 5 10214/[OH2] 1.0 n i and has a value of K 5 1.8 3 10 16 . Note that an equilibrium constant does not formally have units. Nonetheless, the value of the equilibrium constant given assumes that particular units are used for concentration (sometimes referred to a standard states); in this case and in many others, units of molar ity (M) are assumed. With these relations in hand, we can easily calculate the concentration of hydroxide ions in an aqueous solution, given the pH. For example, at pH 5 7.0, we know that [H ] 5 10 7 M and so [OH ] 5 10 14 /10 7 5 10 7 M. In acidic solutions, the concentration of hydrogen ions is higher than 10 7 and, hence, the pH is below 7. For example, in 0.1 M HCl, [H ] 5 10 1 M and so pH 5 1.0 and [OH ] 5 10 14 /10 1 5 10 13 M. Acid–base reactions can disrupt the double helix The reaction that we have been considering between two strands of DNA to and treat it with a the first additions of ate into its 0.8 0.6 0.4 0.2 0 solution of base are made, the component single concentrated base pH rises, but the strands. As the pH 7 8 9 10 11 pH (i.e., with a high concentration of the continues to rise concentration of OH double-helical DNA from 9 to 10, this form a double helix ). As the base is does not change dissociation takes place readily at added, we monitor significantly. becomes essentially pH 7.0. Suppose that the pH and the However, as the pH complete. Why do we take the solution fraction of DNA in approaches 9, the the two strands containing the double-helical form DNA double helix double-helical DNA (Figure 1.16). When begins to dissoci in DNA base pairs to remove certain protons. The FIGURE 1.16 DNA denaturation by the addition of a base. The addition of a base to a solution of double-helical DNA initially at pH 7 causes the double most susceptible proton is the one bound to the N-1 helix to separate nitrogen atom in a guanine base. dissociate? The hydroxide ions can react with bases e s lbu o e l d u c e lo m f o n o it c a r F m r o f l a c il e h- into single strands. The process is half O complete at slightly above pH 9. N NH − O NH N + pKa = 9.7 + H H N N Guanine (G) NH2 NN NH2 Proton dissociation for a substance HA (such as that bound to N-1 on gua nine) has an equilibrium constant defined by the expression Ka 5 [H1][A2]y[HA] The susceptibility of a proton to removal by reaction with a base is often described by its pKa value: pKa 5 2log(Ka) When the pH is equal to the p Ka , we have pH 5 pKa and so 2log[H1] 5 2log([H1][A2]y[HA]) and [H1] 5 [H1][A2]y[HA] Dividing by [H ] reveals that 1 5 [A2]y[HA] and so [A2] 5 [HA] Thus, when the pH equals the p Ka , the concentration of the deprotonated form of the group or molecule is equal to the concentration of the proton ated form; the deprotonation process is halfway to completion. The p Ka for the proton on N-1 of guanine is typically 9.7. When the pH approaches this value, the proton on N-1 is lost (Figure 1.16). Because this proton participates in an important hydrogen bond, its loss substantially destabilizes the DNA double helix. The DNA double helix is also destabi lized by low pH. Below pH 5, some of the hydrogen bond acceptors that participate in base-pairing become protonated. In their protonated forms, these bases can no longer form hydrogen bonds and the double helix sepa rates. Thus, acid–base reactions that remove or donate protons at specific positions on the DNA bases can disrupt the double helix. Buffers regulate pH in organisms and in the laboratory These observations about DNA reveal that a 12 significant change in pH can disrupt molecular structure. The same is true for many other biological the amount of acid added. In contrast, when acid is added to a buffered solu tion, the pH drops more macromolecules; changes in pH can protonate or gradually. Buffers also mitigate the pH increase deprotonate key groups, potentially disrupting structures and initiating harmful reactions. Thus, systems have evolved to mitigate changes in pH in biological systems. Solutions that resist such changes are called buffers . Specifically, when acid is added to an unbuffered aqueous solution, the pH drops in proportion to 10 caused by the addition of base and changes in pH caused by dilution. 8 Compare the result of adding a 1 M solution of the strong acid HCl drop by drop to pure water with adding it to a solution containing 100 mM of the 6 buffer sodium acetate (Na CH 3 COO ; Figure 1.17). The process of 2 drops of acid. However, for the sodium acetate solution, the pH first falls rapidly from its initial value near 10, then changes more gradually until the H p 0 gradually adding known amounts of reagent to a pH reaches 3.5, and then falls more rapidly again. solution with which the 4 Why does the pH decrease so gradually in the reagent reacts while monitoring the results is called a middle of the titration? The answer is that, titration . For pure water, the pH drops from 7 to 15 close to 2 on the addition of the first few 1.3 Chemical Concepts when hydrogen ions are added to this solution, they react with acetate ions to form acetic acid. This reaction consumes some of the added hydrogen ions so that the pH does not drop. Hydrogen ions continue reacting with acetate ions until essentially all of the acetate ion is converted into acetic acid. After this point, added protons remain free in solution and the pH begins to fall sharply again. 0.1 M Na+CH3COO− Gradual pH change Water FIGURE 1.17 Buffer action. The addition of a strong acid, 1 M HCl, to pure water results in an immediate drop in pH to near 2. In contrast, the addition of the acid to a 0.1 M sodium acetate (Na CH3COO ) solution results in a much more gradual change in pH until the pH drops below 3.5. 6050403020100 Number of drops equation to our titration of sodium acetate. The p Ka of acetic acid is 4.75. We can calculate the ratio of the concentration of acetate ion to the concentration of acetic acid as a function of pH by using the Henderson–Hasselbalch equation, slightly rearranged. 16 CHAPTER 1 Biochemistry: An Evolving Science 12 10 60504030201000% We can analyze the effect of the buffer in quantitative terms. The equi librium constant for the deprotonation of an acid is [Acetate ion]y[Acetic acid] 5 [A2]y[HA] 5 10pH2pKa At pH 9, this ratio is 10 9 4.75 5 10 4.25 5 17,800; very little acetic acid 8 percentage has been formed. At pH 4.75 (when Ka 5 [H1][A2]y[HA] 6 the pH equals the p Ka ), the ratio is Taking logarithms of both sides 10 4.75 4.75 5 10 0 5 1. At pH 3, the 4 yields ratio is 10 3 4.75 5 10 1.25 5 0.02; 1 2 2 log(Ka) 5 log([H ]) 1 log([A ]y[HA]) almost all of the acetate ion has been converted into acetic acid. Recalling the definitions of p Ka 0 We can fol and pH and rearranging gives pH low the conversion of acetate ion Number of drops 5 pKa 1 log([A2]y[HA]) into acetic acid over the entire 100% This expression is referred to as titration (Figure 1.18). The graph shows that the region of relatively the Henderson – Hasselbalch constant pH equation . We can apply the From this discussion, we see that a buffer functions FIGURE 1.18 Buffer protonation. When acid is added to sodium acetate, the added hydrogen ions are used to convert acetate ion into acetic acid. Because best close to the p K value of its acid component. a the proton concentration does not increase significantly, the pH remains Physiological pH is typically about 7.4. An important relatively constant until all of the acetate has been converted into acetic acid. buffer in biological systems is based on phosphoric corresponds precisely to the region in which acetate acid (H 3 PO 4 ). The acid can be deprotonated in ion is being protonated to form acetic acid. three steps to form a phosphate ion. Acetic acid H p H H H PO43 H2PO4 HPO42 H3PO4 pKa 2.12 pKa pKa 12.67 7.21 At about pH 7.4, inorganic phosphate exists primarily as a nearly equal mixture of H2PO42 and HPO422. Thus, phosphate solutions function as effective buffers near pH 7.4. The concentration of inorganic phosphate in blood is typically approximately 1 mM, providing a useful buffer against processes that produce either acid or base. We can examine this utility in quantitative terms with the use of the Henderson–Hasselbalch equation. What concentration of acid must be added to change the pH of 1 mM phos phate buffer from 7.4 to 7.3? Without buffer, this change in [H ] corre sponds to a change of 10 7.3 2 10 7.4 M 5 (5.0 3 10 8 2 4.0 3 10 8 ) M 5 1.0 3 10 8 M. Let us now consider what happens to the buffer com ponents. At pH 7.4, [HPO422]y[H2PO42] 5 107.427.21 5 100.19 5 1.55 The total concentration of phosphate, [HPO422] 1 [H2PO42], is 1 mM, Thus, [HPO422] 5 (1.55/2.55) 3 1 mM 5 0.608 mM and [H2PO42] 5 (1/2.55) 3 1 mM 5 0.392 mM GGAGAAGT At pH 7.3, CTGCCGTTACTGCCCTGTGGGGCAAGGTGAACG [HPO42 2]y[H2PO42] 5 107.327.21 5 100.09 5 1.23 TGGA . . . and so is a part of one of the genes that encodes hemoglobin, the oxygen carrier in our blood. This [HPO422] 5 (1.23y2.23) 5 0.552 mM gene is found on the end of chromosome 9 of our 24 and distinct chromosomes. If we were to include the complete sequence of our entire genome, this [H2PO42] 5 (1y2.23) 5 0.448 mM chapter would run to more than 500,000 pages. The 22 Thus, (0.608 2 0.552) 5 0.056 mM HPO4 is sequenc converted into H2PO42, consuming 0.056 mM 5 5.6 ing of our genome is truly a landmark in human 3 10 5 M [H ]. Thus, the buffer increases the history. This sequence contains a vast amount of amount of acid required to produce a drop in pH from information, some of which we can now extract and 7.4 to 7.3 by a factor of 5.6 3 10 5y1.0 3 10 8 5 interpret, but much of which we are only beginning to 5600 compared with pure water. understand. For example, some human diseases have been linked to particular variations in genomic sequence. Sickle-cell anemia, discussed in detail in Chapter 7, is caused by a single base change of an 1.4 The Genomic Revolution Is Transforming A (noted in boldface type in the Biochemistry, Medicine, and Other Fields 17 1.4 The Genomic Revolution Watson and Crick’s discovery of the structure of 18 DNA suggested the hypothesis that hereditary information is stored as a sequence of bases along CHAPTER 1 Biochemistry: An Evolving Science preceding sequence) to a T. We will encounter many long strands of DNA. This remarkable insight other examples of dis eases that have been linked to provided an entirely new way of thinking about specific DNA sequence changes. Determining the biology. However, at the time that it was made, first human genome sequences was a great Watson and Crick’s discovery, though full of challenge. It required the efforts of large teams of potential, remained to be confirmed and many geneticists, molecular biologists, bio chemists, and features needed to be elucidated. How is the computer scientists, as well as billions of dollars, sequence information read and translated into because there was no previous framework for action? What are the sequences of naturally aligning the sequences of various DNA fragments. occurring DNA molecules and how can such sequences be experimentally determined? Through One human genome sequence can serve as a advances in bio chemistry and related sciences, we reference for other sequences. The availability of such reference sequences enables much more now have essentially complete answers to these rapid characterization of partial or complete questions. Indeed, in the past decade or so, genomes from other indi viduals. As we will discuss scientists have deter mined the complete genome sequences of hundreds in Chapter 5, arrays containing millions of target single-stranded DNA molecules with sequences from of different organ isms, including simple microorganisms, plants, animals of varying degrees the reference genome and known or potential of complexity, and human beings. Comparisons of variants are powerful tools. These arrays can these genome sequences, with the use of methods be exposed to mixtures of DNA fragments for a introduced in Chapter 6, have been sources of many particular individual and those single-stranded targets that bind to their complementary strands can insights that have transformed biochemistry. In be determined. This allows many positions within the addition to its experimental and clinical aspects, genome of the indi vidual to be simulaneously biochemistry has now become an information interrogated. science . Genome sequencing has transformed biochemistry and other fields The sequencing of a human genome was a daunting task because it contains approximately 3 billion (3 3 10 9 ) base pairs. For example, the sequence ACATTTGCTTCTGACACAACTGTGTTCACTAGCA ACCTC AAACAGACACCATGGTGCATCTGACTCCTG A Methods for sequencing DNA have also been d e c n e u q e s e m o n e g n a m u h r e p t s o C $100M improv sequencing rate and decreases in cost (Figure 1.19). The availability of such powerful sequencing technology is transforming many $100K fields, including medicine, dentistry, microbiology, pharmacology, and $10K ecology, although a great deal $1K remains to be done to improve the 200620052004200320022001 20112010200920082007 accuracy and precision of the 20132012 interpretation of these large Year genomic and related data sets. ing rapidly, driven by a deep understanding of the biochemistry Each person has a unique sequence of DNA base pairs. How of DNA replication and other different are we from one another processes. This has resulted in both dramatic increases in the DNA at the $1M $10M reveals that, on average, each pair of individuals has a different base in one position per 200 bases; that is, the difference is approximately 0.5%. This variation between individuals who are not closely Human Genome Research Institute. www.genome.gov/sequencingcosts] related is quite substantial compared with genomic level? Examination of genomic variation differences in popu lations. The average difference between two people within one ethnic group is greater than the difference between the averages of two differ ent ethnic groups. The significance of much of this genetic variation is not understood. As noted earlier, variation in a single base within the genome can lead to a disease such as sickle-cell anemia. Scientists have now identified the genetic variations associated with hundreds of diseases for which the cause can be traced to a single gene. For other diseases and traits, we know that variation in many different genes contributes in significant and often complex ways. Many of the most prevalent human ailments such as heart disease are linked to variations in many genes. Furthermore, in most cases, the presence of a particular variation or set of variations does not inevita bly result in the onset of a disease but, instead, leads to a predisposition to the development of the disease. Our own genes are not the only ones that can contribute to health and disease. Our bodies, including our skin, mouth, digestive tract, genito urinary tract, respiratory tract, and other areas, contain large number of microorganisms. These complex communities have been characterized through powerful methods that allow DNA isolated from these biologi cal samples to be sequenced without any previous knowledge of the organisms present. Many of these They appear to play roles in health organisms had not previously been and in diseases such as obesity discovered because they can only and dental caries (Figure 1.20). grow as part of complex In addition to the implications for Nasal communities and thus cannot be understanding human health and isolated through standard microbiological techniques. Remarkably, it appears that we are Gastrointestinal Urogenital outnumbered in our own bodies! Each of us contains approximately Skin ten times more microbial cells than human cells and these microbial cells include many more genes than do our own genomes. These microbiomes differ from site to site, Oral from one person to another and over time in the same individual. disease, the genome sequence is a source of deep of different individuals and populations, we can learn a great deal about human history. On the insight into other aspects of human biology and culture. For example, by comparing the sequences basis of such analysis, a compelling case can be FIGURE 1.19 Decreasing costs of DNA sequencing. Through the Human Genome Project, the cost of DNA sequencing dropped steadily. With the advent of new methods, these costs dropped dramatically and are now approaching $1000 for a complete human genome sequence. [National made that the human species originated in Africa, evolutionary and functional relatives in the genomes and the occurrence and even the timing of of bacteria. Because many studies that are possible important migrations of groups of human beings can in model organisms are difficult or unethical to be dem conduct in human beings, these discoveries have onstrated (Figure 1.21). Finally, comparisons of the many practical implica tions. Comparative genomics has become a powerful science, linking evolu human genome with the genomes of other organisms are confirming the tremendous unity that tion and biochemistry. exists at the level of biochemistry and are revealing 1.20 The human microbiome. Microorganisms cover the human body. key steps in the course of evolution from relatively FIGURE Examination of the microbial communities using DNA sequencing methods simple, single-celled organisms to complex, revealed many previously uncharacterized species. The Venn diagrams multicellular organisms such as human beings. For represent populations of related species as determined by DNA sequence example, many genes that are key to the function of comparisons. The populations present on different body surfaces are largely distinct. [Adapted from www.nature. com/nature/journal/v486/n7402/fig_tab/ the human brain and nervous system have nature11234_F1.html] 50,000–60,000 years ago 46,000–50,000 years ago 15,000–19,000 years ago 150,000 years ago 20,000–30,000 coastal route years ago 15,000 years ago 40,000 years ago 12,500 years ago FIGURE 1.21 Human migrations supported by DNA sequence comparisons. Modern human beings originated in Africa, migrated first to Asia, and then to Europe, Australia, and North and South America. [Adapted from S. Oppenheimer, “Out-of-Africa, the peopling of continents and islands: tracing uniparental gene trees across the map.” Philos. Trans. R. Soc. Lond. B. Biol. Sci. 367(1590):770–784] 20 CHAPTER 1 Biochemistry: An Evolving Science element. Despite the fact that the most important essential dietary factors have been Environmental factors influence human biochemistry Although our genetic makeup (and that of our microbiomes ) is an impor tant factor that contributes to disease susceptibility and to other traits, factors in a person’s environment also are significant. What are these envi ronmental factors? Perhaps the most obvious are chemicals that we eat or are exposed to in some other way. The adage “you are what you eat” has considerable validity; it applies both to substances that we ingest in sig nificant quantities and to those that we ingest in only trace amounts. Throughout our study of biochemistry, we will encounter vitamins and trace elements and their derivatives that play crucial roles in many pro cesses. In many cases, the roles of these chemicals were known for some time, new roles for them continue to first revealed through investigation of deficiency disorders observed in people who do not take in a be discovered. A healthful diet requires a balance of major food sufficient quantity of a particular vitamin or trace Vitamins and minerals 19 FruitsGrains Fats Carbohydrates Dairy groups. In addition to providing vitamins and trace elements, food provides calories in the form FIGURE 1.22 Nutrition. Proper health depends of sub stances that can be Protein broken down to release energy in which food, particularly rich that drives other biochemical foods such as meat, was scarce. processes. Proteins, fats, and With the development of carbohydrates provide the agriculture and modern building blocks used to construct economies, rich foods are now the molecules of life (Figure plentiful in parts of the world. 1.22). Finally, it is possible to get Some of the most prevalent too much of a good thing. diseases in the so-called Human beings evolved under developed world, such as circumstances Vegetables Protein Just as vitamin deficiencies and genetic diseases have revealed funda mental principles of biochemistry and biology, investigations of variations choosemyplate.gov] in behavior and their linkage to genetic and heart disease and diabetes, can be attributed to the biochemical factors are potential sources of great large quantities of fats and carbohydrates present in insight into mechanisms within the brain. For modern diets. We are now developing a deeper example, studies of drug addiction have revealed understanding of the biochemical consequences of neural circuits and biochemical pathways that these diets and the inter play between diet and greatly influence aspects of behavior. Unraveling the genetic factors. inter play between biology and behavior is one of the Chemicals are only one important class of great challenges in modern science, and environmental factors. Our behaviors also have biochemistry is providing some of the most important biochemical consequences. Through physical activ concepts and tools for this endeavor. ity, we consume the calories that we take in, Genome sequences encode proteins and patterns of expression ensuring an appropriate bal ance between food The structure of DNA revealed how information is intake and energy expenditure. Activities ranging from exercise to emotional responses such as fear stored in the base sequence along a DNA strand. But what information is stored and how is it and love may activate specific biochemical expressed? The most fundamental role of DNA is to pathways, leading to changes in levels of gene encode the sequences of proteins. Like DNA, expression, the release of hormones, and other consequences. Furthermore, the interplay between proteins are linear polymers. However, proteins differ from DNA in two important ways. First, proteins biochemistry and behavior is bidirectional. Just as our biochemistry is affected by our behavior, so, too, are built from 20 building blocks, called amino acids, rather than just four, as in DNA. The chemical our behavior is affected, although certainly not completely determined, by our genetic makeup and complex ity provided by this variety of build other aspects of our biochemistry. Genetic factors associated with a range of behavioral characteristics 21 1.4 The Genomic Revolution have been at least tentatively identified. a solution of double-helical molecules. A similar spontaneous folding process gives proteins their three-dimensional structure. A bal ance of hydrogen bonding, van der on an appropriate combination of food groups (fruits, vegetables, protein, grains, dairy) (left) to provide an optimal mix of biochemicals (carbohydrates, proteins, fats, vitamins, and minerals) (right). [Adapted from www. ing blocks enables proteins 1 2 3 to per form a wide range of Amino acid sequence 1 functions. Second, proteins spontaneously fold into elaborate three-dimensional structures, determined only by their amino acid sequences (Figure 1.23). We have explored in depth 123 how solutions containing two appropri ate strands of Amino acid sequence 2 DNA come together to form Waals interactions, and hydrophobic interactions overcomes the entropy lost in going from an The fundamental unit of hereditary information, the gene, is becom ing increasingly difficult to precisely define as our knowledge of the com plexities of genetics and genomics increases. The genes that are simplest to define encode the sequences of proteins. For these protein-encoding genes, a block of DNA bases encodes the amino acid sequence of a spe cific protein molecule. A set of three bases along the DNA strand, called a codon, determines the identity of one amino acid within the protein sequence. The set of rules that links the DNA sequence to the encoded protein sequence is called the genetic code . One of the biggest surprises from the sequencing of the human genome is the small number of pro tein-encoding genes. Before the genome-sequencing project began, the consensus view was that the human genome would include approximately 100,000 protein-encoding genes. The current analysis suggests that the actual number is between 20,000 and 25,000. We shall use an estimate of 21,000 throughout this book. However, additional mechanisms allow many genes to encode more than one protein. For example, the genetic information in some genes is translated in more than one way to produce a set of proteins that differ from one another in parts of their amino acid sequences. In other cases, proteins are modified after they have been syn thesized through the addition of accessory chemical groups. Through these indirect mechanisms, much more complexity is encoded in our genomes than would be expected from the number of protein-encoding genes alone. On the basis of current knowledge, the proteinencoding regions account for only about 3% of the human genome. What is the function of the rest of the DNA? Some of it contains information that regulates the expression of specific genes (i.e., the production of specific proteins) in particular cell types and physiological conditions. Essentially every human FIGURE 1.23 Protein folding. Proteins are linear polymers of amino acids that fold into elaborate structures. The sequence of amino acids determines the three-dimensional structure. Thus, amino acid sequence 1 gives rise only to a protein with the shape depicted in blue, not the shape depicted in red. 22 CHAPTER 1 Biochemistry: An Evolving Science cell contains the same DNA genome, yet cell types differ considerably in the proteins that they produce. For example, hemoglobin is expressed only in precursors of red blood cells, even though the genes for hemoglobin are present in essentially every cell. Specific sets of genes are expressed in response to hormones, even though these genes are not expressed in the same cell in the absence of the hormones. The control regions that regulate such differences account for only a small amount of the remainder of our genomes. The truth is that we do not yet understand all of the function of much of the remainder of the DNA. Some of it is sometimes referred to as “junk”—stretches of DNA that were inserted at some stage of evolution and have remained. In some cases, this DNA may, in fact, serve important functions. In others, it may serve no function but, because it does not cause significant harm, it has remained. APPENDIX: Visualizing Molecular Structures I: Small Molecules The authors of a biochemistry textbook face the problem of trying to present threedimensional molecules in the two Z X C W W ZW of biomolecules dimensions available on the printed page. The ZX≡≡ interplay between the three-dimensional structures Y X frequently and their biological functions will be Y discussed extensively throughout Fischer this book. Toward this end, we will projection use representations that, although of necessity are rendered in two dimensions, emphasize the threedimensional struc tures of molecules. Stereochemical Renderings Most of the chemical formulas in this book are drawn to depict the geometric arrangement of atoms, crucial to chemical bonding and reactivity, as accurately as possi ble. For example, the carbon atom of methane is tetra hedral, with H–C–H angles of 109.5 degrees, whereas the carbon atom in formaldehyde has bond angles of 120 degrees. Y Stereochemical rendering carbon are represented by horizontal and vertical lines from the substituent atoms to the carbon atom, which is assumed to be at the center of the cross. By convention, the horizontal bonds are assumed to project out of the page toward the viewer, whereas the vertical bonds are assumed to project behind the page away from the viewer. Molecular Models for Small Molecules For depicting the molecular architecture of small molecules in more detail, two types of models will often be used: space filling and ball and stick. These models show structures at In a Fischer projection, the bonds to the central HH C H H C H H the atomic level. O Methane Formaldehyde space-filling models are the most realistic. The size and position of an atom in a space 1. Space-Filling Models . The shown in Figure 1.24. To illustrate the correct stereochemistry about tetra hedral carbon atoms, wedges will be used to depict the direction of a bond into or out of the plane of the page. A solid wedge with the broad end away from the carbon atom denotes a bond coming toward the viewer out of the plane. A dashed wedge, with its broad end at the carbon atom, represents a bond going away from the viewer behind the plane of the page. The remaining two bonds are depicted as straight lines. 2. Ball-and-Stick Models . Ball-and-stick models are not as realistic as space-filling models, because the atoms are depicted as spheres of radii smaller than their van der Waals radii. However, the bonding arrangement is easier to see because the bonds are explicitly represented as sticks. In an illustration, the taper of a stick, representing parallax, tells which of a pair of bonded atoms is closer to the reader. A ball-and-stick Fischer Projections model reveals a complex structure more clearly than a Although representative of the actual structure of a com pound, stereochemical structures are often difficult space filling model does. Ball-and-stick models of several simple molecules are shown in Figure 1.24. to draw quickly. An alternative, less-representative 23 method of depicting structures with tetrahedral carbon Problems centers relies on the use of Fischer projections . filling model are determined by its bonding properties and van der Waals radius, or contact distance. A van der Molecular models for depicting large molecules will be Waals radius describes how closely two atoms can discussed in the appendix to Chapter 2. approach each other when they are not linked by a covalent bond. The colors of the model are set by Water Acetate Formamide Cysteine SH convention. Carbon, black Hydrogen, white Nitrogen, blue Oxygen, red Sulfur, yellow Phosphorus, purple Space-filling models of several simple molecules are FIGURE 1.24 Molecular representations. Structural formulas (bottom), ball-and- molecules are shown. Black 5 carbon, red 5 oxygen, white 5 hydrogen, yellow 5 sulfur, blue 5 nitrogen. stick models (top), and space-filling representations (middle) of selected +H N 3 O H3C C H − H2N C H O H2O O O OC− KEY TERMS biological macromolecule (p . 2) metabolite (p. 2) deoxyribonucleic acid (DNA) (p. 2) protein (p. 2) Eukarya (p. 3) Bacteria (p. 3) Archaea (p. 3) eukaryote (p. 3) prokaryote (p. 3) double helix (p. 5) covalent bond (p. 5) resonance structure (p. 7) ionic interaction (p. 7) hydrogen bond (p. 8) van der Waals interaction (p. 8) hydrophobic effect (p. 9) hydrophobic interaction (p. 9) entropy (p. 11) enthalpy (p. 11) free energy (Gibbs free energy) (p. 12) pH (p. 13) p Ka value (p. 14) buffer (p. 15) predisposition (p. 18) microbiome (p. 19) amino acid (p. 21) genetic code (p. 21) PROBLEMS 1. Donors and acceptors . Identify the hydrogen-bond donors and acceptors in each of the four bases on page 4. 2. Resonance structures . The structure of an amino acid, tyro sine, is shown here. Draw an alternative resonance structure. O H H H H H H CH2 C +H N COO− 3 24 CHAPTER 1 Biochemistry: An Evolving Science (b) DH 5 2 84 kJ mol 1 ( 2 20 kcal mol 1 ), 3. It takes all types . What types of noncovalent bonds hold together the following solids? ( a ) Table salt ( NaCl ), which contains Na and Cl ions. (b) Graphite (C), which consists of sheets of covalently bonded carbon atoms. 4. Don’t break the law . Given the following values for the changes in enthalpy ( DH ) and entropy ( DS ), which of the following processes can take place at 298 K without violat ing the Second Law of Thermodynamics? (a) DH 5 2 84 kJ mol 1 ( 2 20 kcal mol 1 ) , DS 5 1 125 J mol 1 K 1 ( 1 30 cal mol 1 K 1) DS 5 2 125 J mol 1 K 1 ( 2 30 cal mol 1 K 1 ) (c) DH 5 1 84 kJ mol 1 ( 1 20 kcal mol 1 ), DS 5 1 125 J mol 1 K 1 ( 1 30 cal mol 1 K 1 ) (d) DH 5 1 84 kJ mol 1 ( 1 20 kcal mol 1 ), DS 5 2 125 J mol 1 K 1 ( 2 30 cal mol 1 K 1 ) 5. Double-helix-formation entropy . For double-helix forma tion, D G can be measured to be 2 54 kJ mol 1 ( 2 13 kcal mol 1 ) at pH 7.0 in 1 M NaCl at 25 8 C (298 K). The heat released indicates an enthalpy change of 2 251 kJ mol 1 ( 2 60 kcal mol 1 ) . For this process, calculate the entropy change for the system and the entropy change for the surroundings. 6. Find the pH . What are the pH values for the following solutions? (a) 0.1 M HCl (b) 0.1 M NaOH (c) 0.05 M HCl (d) 0.05 M NaOH 7. A weak acid . What is the pH of a 0.1 M solution of acetic acid ( p Ka 5 4. 75 ) ? (Hint: Let x be the concentration of H ions released from acetic acid when it dissociates. The solutions to a quadratic equation of the form ax2 1 bx 1 c = 0 are x 5 (2b 6 2b2 2 4ac)y2a.) 8. Substituent effects . What is the pH of a 0.1 M solution of chloroacetic acid ( ClCH 2 COOH, pKa 5 2 . 86)? 9. Water in water . Given a density of 1 g/ml and a molecu lar weight of 18 g/mol, calculate the concentration of water in water. 10. Basic fact . What is the pH of a 0.1 M solution of ethylamine, given that the p Ka of ethylammonium ion ( CH 3 CH 2 NH 3+ ) is 10.70? 11. Comparison . A solution is prepared by adding 0.01 M acetic acid and 0.01 M ethylamine to water and adjusting the pH to 7.4. What is the ratio of acetate to acetic acid? What is the ratio of ethylamine to ethylammonium ion? 12. Concentrate . Acetic acid is added to water until the pH value reaches 4.0. What is the total concentration of the added acetic acid? 13. Dilution . 100 mL of a solution of hydrochloric acid with pH 5.0 is diluted to 1 L . What is the pH of the diluted solution? 14. Buffer dilution . 100 mL of a 0.1 mM buffer solution made from acetic acid and sodium acetate with pH 5.0 is diluted to 1 L . What is the pH of the diluted solution? 17. What’s the ratio? An acid with a p Ka of 8.0 is present in a solution with a pH of 6.0. What is the ratio of the proton ated to the deprotonated form of the acid? 18. Phosphate buffer . What is the ratio of the concentra tions of H2PO4 and HPO42 at (a) pH 7.0; (b) pH 7.5; (c) pH 8.0? 19. Neutralization of phosphate . Given that phosphoric acid (H 3 PO 4 ) can give up three protons with different pK a values, sketch a plot of pH as a function of added drops of sodium hydroxide solution, starting with a solution of phosphoric acid at pH 1.0. 20. Buffer capacity . Two solutions of sodium acetate are prepared, one with a concentration of 0.1 M and the other with a concentration of 0.01 M . Calculate the pH values when the following concentrations of HCl have been added to each of these solutions: 0.0025 M , 0.005 M , 0.01 M , and 0.05 M . 21. Buffer preparation . You wish to prepare a buffer con sisting of acetic acid and sodium acetate with a total acetic acid plus acetate concentration of 250 mM and a pH of 5.0. What concentrations of acetic acid and sodium acetate should you use? Assuming you wish to make 2 liters of this buffer, how many moles of acetic acid and sodium acetate will you need? How many grams of each will you need (molecular weights: acetic acid 60 . 05 g mol 1 , sodium acetate, 82 . 03 g mol 1 )? 22. An alternative approach . When you go to prepare the buffer described in Problem 21, you discover that your laboratory is out of sodium acetate, but you do have sodium hydroxide. How much (in moles and grams) acetic acid and sodium hydroxide do you need to make the buffer? 23. Another alternative . Your friend from another labora tory was out of acetic acid, so tries to prepare the buffer in Problem 21 by dissolving 41.02 g of sodium acetate in water, carefully adding 180.0 ml of 1 M HCl , and adding more water to reach a total volume of 2 liters . What is the total concentration of acetate plus acetic acid in the solu tion? Will this solution have pH 5.0? Will it be identical with the desired buffer? If not, how will it differ? 24. Blood substitute . As noted in this chapter, blood con tains a total concentration of phosphate of approximately 1 mM and typically has a pH of 7.4. You wish to make 100 liters of phosphate buffer with a pH of 7.4 from NaH 2 PO 4 (molecular weight, 119 . 98 g (molec mol 1 ) and Na 2 HPO 4 ular weight, 141 . 96 g 1 mol ) . How much of each (in grams) do you need? 15. Find the pKa . For an acid HA, the concentrations of HA and A are 0.075 and 0.025, respectively, at pH 6.0. 25. A potential problem . You wish to make a buffer with pH 7.0. You combine 0.060 grams of acetic acid What is the p Ka value for HA? 16. pH indicator . A dye that is an acid and that appears and 14.59 grams of sodium acetate and add water to as different colors in its protonated and deprotonated yield a total vol ume of 1 liter . What is the pH? Will this forms can be used as a pH indicator. Suppose that you be the useful pH 7.0 buffer you seek? have a 0.001 M solution of a dye with a p Ka of 7.2. From 26. Charge! Suppose two phosphate groups in DNA the color, the concentration of the protonated form is (each with a charge of 2 1) are separated by 12 Å. found to be 0.0002 M . Assume that the remainder of What is the energy of the ionic interaction between the dye is in the deprotonated form. What is the pH of these two phos phates assuming a dielectric constant of 80? Repeat the calculation assuming a dielectric the solution? constant of 2. 27. Vive la différence . On average, estimate how many base differences there are between two human beings. 25 Problems 28. Epigenomics . The human body contains many distinct cell types yet almost all human cells contain the same genome with 21,000 genes. The distinct cell types are pri marily due to differences in gene expression. Assume that one set of 1000 genes is expressed in all cell types and that the remaining 20,000 genes can be divided into sets of 1000 genes that are either all expressed or all not expressed in a given cell type. How many different cell types are possible if each cell type expresses 10 sets of these genes? Note that the number of combinations of n objects into m sets is given by n !/ (m!(n-m)!) where n! = 1*2* …*(n 2 1)*n. 29. Predispositions in populations . Assume that 10% of the members of a population will get a particular disease over the course of their lifetime. Genomic studies reveal that 5% of the population have sequences in their genomes such that their probability of getting the disease over the course of their lifetimes is 50%. What is the average lifetime risk of this disease for the remaining 95% of the population with out these sequences? Protein Composition and Structure N Leu Tyr Gln Leu Glu Asn Tyr C 2 within this sequence can fold into regular structures (the secondary structure), such as the a-helix. Entire chains fold into well-defined structures (the tertiary structure)—in this case, a single insulin molecule. Such structures assemble with other chains to form arrays such as the complex of six insulin molecules shown at the far right (the quaternary structure). These arrays can often be induced to form well-defined crystals (photograph at left), which allows a determination of these structures in detail. [Photograph from Christo Nanev.] CHAPTER Leu Insulin is a protein hormone, crucial for maintaining blood sugar at appropriate levels. (Below) Chains of amino acids in a specific sequence (the primary Glu structure) define a protein such as insulin. Amino acids close to one another Secondary structure Tertiary structure Quarternary structure Primary structure dimensional structure formed by hydrogen bonds between amino acids near one another is called secondary structure, whereas tertiary structure is roteins are the most versatile macromolecules formed by long-range interactions between amino acids. Protein function depends directly on this three dimensional structure (Figure 2.1). Thus, proteins are in living systems and the embodiment of the transition from the oneserve crucial functions in essentially all biological dimensional world of sequences to the threeprocesses. They func tion as catalysts, transport and dimensional world of molecules capable of diverse store other molecules such as oxygen, pro vide activities . Many proteins also display mechanical support and immune protection, generate O U T L I N E movement, transmit nerve impulses, and control growth and differentiation. Indeed, much of this book 2.1 Proteins Are Built from a Repertoire of 20 Amino Acids will focus on understanding what proteins do and 2.2 Primary Structure: Amino Acids Are Linked by Peptide Bonds to how they perform these functions. Form Polypeptide Chains Several key properties enable proteins to participate 2.3 Secondary Structure: Polypeptide Chains Can Fold into Regular in a wide range of functions. P Structures Such As the Alpha Helix, the Beta Sheet, and Turns and Loops 1. Proteins are linear polymers built of monomer units called amino acids, which are linked end to end. 2.4 Tertiary Structure: Water-Soluble Proteins Fold into Compact The sequence of linked amino acids is called the Structures with Nonpolar Cores primary structure. Remarkably, proteins 2.5 Quaternary Structure: Polypeptide Chains Can Assemble into spontaneously fold up into three-dimensional Multisubunit Structures structures that are determined by the sequence of 2.6 The Amino Acid Sequence of a Protein Determines Its Three amino acids in the protein polymer. ThreeDimensional Structure 27 DNA FIGURE 2.1 Structure dictates function. A protein component of the DNA 2. Proteins contain a wide range of functional groups . These functional groups include alcohols, thiols, thioethers, carbox ylic acids, carboxamides, and a variety of basic groups. Most of these groups are chemically reactive. When combined in various sequences, this array of functional groups accounts for the broad spectrum of protein function. For instance, their reactive properties are essential to the function of enzymes, the proteins that catalyze specific chemical reactions in bio logical systems (Chapters 8 through 10). 3. Proteins can interact with one another and with other bio logical macromolecules to form complex assemblies . The proteins within these assemblies can act synergistically to generate capabilities that individual proteins may lack. Examples of these assemblies include macromolecular machines that repli cate DNA, transmit signals within cells, and enable muscle cells to contract (Figure 2.2). replication machinery surrounds a section quaternary structure, in which the functional protein is (A) composed of several distinct polypeptide chains. of DNA double helix depicted as a cylinder. The protein, which consists of two identical subunits (shown in red and yellow), acts as a clamp that allows large segments of DNA to be copied without the replication machinery dissociating from the DNA. Myofibrils [Drawn from 2POL.pdb.] Single muscle fiber (cell) Plasma membrane Single myofibril Nucleus Sarcomere I band A band I band Z line (B) (C) FIGURE 2.2 A complex protein assembly. (A) A single muscle cell contains multiple myofibrils, each Thick of which is filaments comprised of numerous repeats Z line of a complex protein assembly known as the sarcomere. (B) The banding pattern of a sarcomere, evident by electron microscopy, is caused by (C) the interdigitation of filaments made up of many individual proteins. [(B) Courtesy of Dr. Hugh Huxley.] H zone Thin filaments 4. Some proteins are quite rigid, whereas others display considerable flexibility . Rigid units can function as structural elements in the cytoskeleton (the internal scaffolding within cells) or in connective tissue. Proteins with some flexibility may act as hinges, springs, or levers. In addition, conformational changes within proteins enable the regulated assembly of larger protein complexes as well as the transmission of information within and between cells Iron 2.1 Proteins Are Built from a Repertoire of 20 Amino Acids Amino acids are the building blocks of proteins. An amino acid consists of a central carbon atom, called the carbon, linked to an amino group, a carboxylic acid group, a hydrogen atom, and a distinctive R group. The R group is often referred to as the side chain . With four different groups connected to the tetrahe dral a -carbon atom, a -amino acids are chiral: they may exist in one or the other of two mirror-image forms, called the L isomer and the D isomer (Figure 2.4). Only L amino acids are constituents of proteins . For almost all amino acids, the L isomer has S (rather than R ) absolute configuration (Figure 2.5). What is the basis for the preference for L amino acids? The answer has been lost to evolutionary history. It is possible that the preference for L over D amino acids was a consequence of a chance selection. However, there is evidence that L amino acids are slightly more soluble than a racemic mixture of D and L amino 29 2.1 Amino Acids FIGURE 2.3 Flexibility and function. On binding iron, the protein lactoferrin undergoes a substantial change in conformation that allows other molecules to distinguish between the iron-free and the iron-bound forms. [Drawn from 1LFH.pdb and 1LFG.pdb.] Notation for distinguishing stereoisomers The four different substituents of an asymmetric carbon atom are assigned a priority according to atomic number. The lowest-priority substituent, often hydrogen, is pointed away from the viewer. The configuration about the carbon atom is called S (from the Latin sinister, “left”) if the progression from the highest to the lowest priority is counterclockwise. The configuration is called R (from the Latin rectus, “right”) if the progression is acids, which tend to form crystals. This small solubility clockwise. difference could have been (also called zwitterions ). In the amplified over time so that the L dipolar form, the amino group is isomer became dominant in protonated solution. Amino acids in R solution at neutral pH exist predominantly as dipolar ions H (4) (3) (1) (2) Cα HRRHC C α COO− NH3+ α (}NH3 ) and the carboxyl group is deprotonated (}COO ). The ionization state of an amino acid varies with pH (Figure 2.6). In acid solution (e.g., pH 1), the amino group is protonated (}NH3 ) and the carboxyl group is not dissociated (}COOH). As the pH is raised, the carboxylic acid is the first group to give up a proton, inasmuch as its p Ka is near 2. The dipolar form persists until the pH approaches 9, when the protonated amino group loses a proton. NH3+ NH3+ COO− COO− L isomer D isomer FIGURE 2.4 The L and D isomers of amino acids. The letter R refers to the side chain. The L and D isomers are mirror images of each other. FIGURE 2.5 Only L amino acids are found in proteins. Almost all L amino acids have an S absolute configuration. The counterclockwise direction of the arrow from highest- to lowest-priority substituents indicates that the chiral center is of the S configuration. 30 CHAPTER 2 Protein Composition and Structure RH FIGURE 2.6 Ionization state as a function of pH. The acids is altered by a change in R HC ionization state of amino pH. H+ H+ The zwitterionic form predominates near physiological pH. RHC C H + COOH + +H N H N COO– H N COO– H+ 3 3 2 o it a r t n e c n o C Zwitterionic form Both groups protonated Both groups deprotonated n 0 2 4 6810 12 14 pH Twenty kinds of side chains varying in size, shape, charge, hydrogen bonding capacity, hydrophobic character, and chemical reactivity are com monly found in proteins. Indeed, all proteins in all species—bacterial, archaeal, and eukaryotic—are constructed from the same set of 20 amino acids with only a few exceptions. This fundamental alphabet for the con struction of proteins is several billion years old. The remarkable range of functions mediated by proteins results from the diversity and versatility of these 20 building blocks. Understanding how this alphabet is used to create the intricate three-dimensional structures that enable proteins to carry out so many biological processes is an exciting area of biochemistry and one that we will return to in Section 2.6. Although there are many ways to classify amino acids, we will sort these molecules into four groups, on the basis of the general chemical characteris tics of their R groups: 1. Hydrophobic amino acids with nonpolar R groups 2. Polar amino acids with neutral R groups but the charge is not evenly distributed 3. Positively charged amino acids with R groups that have a positive charge at physiological pH 4. Negatively charged amino acids with R groups that have a negative charge at physiological pH Hydrophobic amino acids. The simplest amino acid is glycine, which Glycine (Gly, G) Alanine (Ala, A) has a single hydrogen atom as its side chain. With two hydrogen atoms bonded to the a -carbon atom, glycine is unique in being achiral . Alanine, the next simplest amino acid, has a methyl group (}CH 3 ) as its side chain (Figure 2.7). C H H H H CH2 C CH3 H3C CH3 Valine (Val, V) Proline (Pro, P) 2C Leucine (Leu, L) CH3 H +H N 3 COO– +H N 3 COO– C H2C H2COO– H 2C CH3 C H + COO– H N C 3 +H N 3 (Gly, G) (Ala, A) H2 H2C CH2 N+ COO– Alanine C + COO– H N C 3 H Proline (Pro, P) CH3 CH2 +H N 3 C +H N 3 CH3 COO– H Glycine HC CH COO– H CH H N+ COO– CH3 H CH3 C Valine (Val, V) H CH3 C H CH2 H +H N 3 Leucine (Leu, L) COO– C Isoleucine (Ile, I) H3C CH3 H H C CH * 2 H 3 Methionine (Met, M) S H2C Tryptophan (Trp, W) H H H Phenylalanine (Phe, F) C +H N 3 C H COO– +H N 3 C +H N 3 CH2 H C CH2 H H HNH COO– CH2 H +H N COO– 3 CH3 C H CH3 CH2 CH2 H C CH3 S +H N COO– HC 3 COO– +H N 3 Isoleucine (Ile, I) CH HC C CH2 HC CH COO– HN HC CH HC CH2 C CH H H H CH CC + COO– H N C 3 Phenylalanine CH H H CH2 C (Phe, F) representation (see the Appendix to Chapter 1). FIGURE 2.7 Structures of hydrophobic amino acids. For each amino acid, a ball-and-stick model The additional chiral center in isoleucine is indicated (top) shows the arrangement of atoms and bonds in by an asterisk. space. A stereochemically realistic formula (middle) + COO– H3N C shows the geometric arrangement of bonds around atoms, and a Fischer projection (bottom) shows all H Tryptophan bonds as being perpendicular for a simplified 31 Methionine (Met, M) (Trp, W) Tyrosine (Tyr, Y) Asparagine (Asn, N) Glutamine (Gln, Q) O HO CH H Threonine (Thr, T) Serine (Ser, S) CH3 H COO– H C * H2N OH H H C NH2 +H N 3 +H COO– OC O CH H 2 3N H OH H C H COO– H (Ser, S) Threonine (Thr, T) C HC CH (Cys, C) C CH +H N 3 HCH2 +H N COO– 3 O NH2 C O NH2 C HC C CH2 2 +H N COO– 3 C H Serine HO CH3 C + COO– H N 3 Cysteine H CH +H N COO– 3 OH CH2 +H N 3 H2C H H Asparagine CH2 CH2 + COO– H N C 3 CH2 + COO– H N C 3 Tyrosine (Tyr, Y) CH COO– (Asn, N) Glutamine (Gln, Q) H SH HCH2 +H N COO– 3 SH CH2 + COO– H N 3 C H Cysteine (Cys, C) FIGURE 2.8 Structures of the polar amino acids. The additional chiral center in threonine is indicated by an asterisk. HN tend to cluster together rather than contact water. The three-dimensional structures of water soluble proteins are stabilized by this tendency of hydrophobic groups to come together, which is called the hydrophobic effect (p. 9). The different sizes and shapes of these hydrocarbon side chains enable them to pack together to form compact structures with little empty space. Proline also has an aliphatic side chain, but it differs from other members of the set of 20 in that its side chain is bonded to both the nitrogen and the a -carbon atoms, yielding a pyrrolidine ring. Proline markedly influences protein architecture because its cyclic struc ture makes it more conformationally restricted than the other amino acids. Two amino acids with relatively simple aromatic side chains are part of the fundamental repertoire. Phenylalanine, as its name indicates, contains a phe nyl ring attached in place of one of the hydrogen atoms of alanine. Tryptophan has an indole group joined to a methylene (}CH 2}) group; the indole group comprises two fused rings containing an NH group. Phenylalanine is purely hydrophobic, whereas tryptophan is less so because of its NH group. Larger hydrocarbon side chains are found in valine, leucine, and isoleucine . Methionine contains a Polar amino acids. Six amino acids are polar but largely aliphatic side chain that includes a thioether uncharged. Three amino acids, serine, threonine, (}S}) group. The side chain of isoleucine includes an and tyrosine, contain hydroxyl groups (}OH) attached additional chiral center; only the isomer shown in to a hydrophobic side chain (Figure 2.8). Serine can Figure 2.7 is found in proteins. The larger aliphatic be thought of as a version side chains are especially hydrophobic; that is, they of alanine with a hydroxyl group attached, threonine resembles valine with a Indole 32 hydroxyl isoleucine, group in place con tains an of one of additional valine’s asymmetric methyl center; again, groups, and only one tyrosine is a isomer is version of present in phenylalanine proteins. with the In addition, hydroxyl the set group includes replacing a asparagine NH3 hydrogen and glutamine H2C atom on the , two amino + aromatic ring.acids that con Arginine (Arg, R) The hydroxyl tain a terminal group makes carboxamide . these amino The side acids much chain of more glutamine is hydrophilic one (water loving) methylene and reactive group longer than their than that of hydrophobic asparagine. Lysine analogs. H2N (Lys, K) Threonine, like H N + C HN NH2 H C Histidine (His, H) Cysteine is structurally similar to ser ine but contains a sulfhydryl, or thiol (}SH), group in place of the hydroxyl (}OH) group. The sulfhydryl group is H2C H2C H C CH2 CH2 H much more reactive. Pairs of sulfhydryl + C CH2 CH2 N H C H C particularly impor tant in stabilizing some proteins, NH3+ as will be discussed shortly. CH2 NH CH2 +H N COO– 3 Positively charged amino CH2 +H N COO– acids. We turn 3 groups may come together to form disul fide H3N COO– bonds, which are + H2N NH2 C H N HC now to amino acids hydrophilic. Lysine CH CH 2 2 and arginine have with complete posi N long CH tive charges that CH CH 2 2 render them highly C H amino group and arginine side chains that terminate by a guanidinium group. Lysine with groups that are (Lys, K) Histidine contains an positively charged at + COO– H N C 3 imidazole neutral pH. Lysine is + COO– H N C H 3 capped by a primary CH2 Arginine (Arg, R) + COO– H N C 3 H Histidine (His, H) group, an aromatic ring that also can be positively charged (Figure 2.9). (Figure 2.10). Histidine is often found in the active value near 6, the imidazole group can sites of enzymes, where NH2 With a p Ka the imidazole ring can bind and release protons in be uncharged or posi tively charged near neutral the course of enzymatic + N pH, depending on its local environment HH FIGURE 2.9 Positively charged amino acids lysine, arginine, and histidine. reactions. C H2N NH2 C NC C H H are charged derivatives of called aspartate and glutamate to This set of amino acids contains asparagine and glutamine (Figure emphasize that, at physiological two with acidic side chains: 2.8), with a carboxylic acid in Guanidinium aspartic acid and glutamic acid place of a carboxamide. Aspartic Imidazole (Figure 2.11). These amino acids acid and glutamic acid are often present in the acid form Negatively charged amino acids. HH pH, their side chains usually lack a proton that is and hence are negatively charged. Nonetheless, these side chains can accept N N HC HC protons in some proteins, often with functionally important consequences. + CH CH Seven of the 20 amino acids have readily ionizable side chains. These H+ N N C C H 7 amino acids are able to donate or accept protons to facilitate reactions as CH2 CH2 well as to form ionic bonds. Table 2.1 gives equilibria and typical p Ka val ues for ionization of the side chains of tyrosine, cysteine, arginine, lysine, H H H+ C bind or release protons near physiological pH. C NCOH NCOH FIGURE 2.10 Histidine ionization. Histidine can 33 Aspartate (Asp, D) Glutamate (Glu, E) TABLE 2.1 Typical pKa values of ionizable groups in proteins Group Acid Base Typical pKa* O O CHO O O Aspartic acid Terminal a-carboxyl group 3.1 C–O C– Glutamic acid 4.1 C H O O– – CO O N O H O C Histidine 6.0 C H2 C H CH2 +H N COO– 3 +H N COO– 3 + C CH2 H N + N H H NH N H Terminal a-amino group 8.0 H OO– C H OO– N H H Cysteine 8.3 CH2 C CH2 S– S CH2 H Tyrosine 10.9 +H COO– 3N C H Glutamate H Aspartate +H COO– O– N O H 3N C NH + (Glu, E) Lysine 10.8 H (Asp, D) H H FIGURE 2.11 Negatively charged H H amino acids. N N H + H H Arginine 12.5 N H C NH N H C NH *pKa values depend on temperature, ionic strength, and the microenvironment of the ionizable group histidine, and aspartic and glutamic acids in proteins. Two other groups in proteins—the terminal a -amino group and the terminal a carboxyl group—can be ionized, and typical p Ka values for these groups also are included in Table 2.1. Amino acids are often designated by either a three-letter abbreviation or a one-letter symbol (Table 2.2). The abbreviations for amino acids are the first three letters of their names, except for asparagine (Asn), glutamine (Gln), isoleucine (Ile), and tryptophan (Trp). The symbols for many amino acids are the first letters of their names (e.g., G for glycine and L for leucine); TABLE 2.2 Abbreviations for amino acids Three-letter One-letter abbreviation abbreviation Amino acid Amino acid Three-letter abbreviation One-letter abbreviation Alanine Ala A Methionine Met M Arginine Arg R Phenylalanine Phe F Asparagine Asn N Proline Pro P Aspartic acid Asp D Serine Ser S Cysteine Cys C Threonine Thr T Glutamine Gln Q Tryptophan Trp W Glutamic acid Glu E Tyrosine Tyr Y Glycine Gly G Valine Val V Histidine His H Asparagine or Isoleucine Ile I aspartic acid Asx B Leucine Leu L Glutamine or Lysine Lys K glutamic acid Glx Z 34 H H C C NH 2C C O H 2C H OH N H X 2C NH C C 2 CH 2 CO FIGURE 2.12 Undesirable C cyclization of serine would form a strained, four-membered ring 2 Homoserine H H O HX+ HOH O Serine X H CHO form a stable, five-membered ring, potentially resulting in peptide-bond cleavage. The C N H C O HX+ and is thus disfavored. X can be an amino group from a neighboring amino acid or another potential leaving group. the other symbols have been agreed on by convention. These abbreviations and symbols are an integral part of the vocabulary of biochemists. How did this particular set of amino acids become the building blocks of proteins? First, as a set, they are diverse: their structural and chemical properties span a wide range, endowing proteins with the versatility to assume many functional roles. Second, many of these amino acids were probably available from prebiotic reactions; that is, from reac tions that took place before the origin of life. Finally, other possible amino acids may have simply been too reactive. For example, amino acids such as homoserine and homocysteine tend to form five-membered cyclic forms that limit their use in proteins; the alternative amino acids that are found in proteins—serine and cysteine—do not readily cyclize, because the rings in their cyclic forms are too small (Figure 2.12). 2.2 Primary Structure: Amino Acids Are Linked by Peptide Bonds to Form Polypeptide Chains Proteins are linear polymers formed by linking the a -carboxyl group of one amino acid to the a -amino group of another amino acid. This type of linkage is called a peptide bond or an amide bond . The formation of a dipeptide from two amino acids is accompanied by the loss of a water molecule (Figure 2.13). The equilibrium of this reaction lies on the side of hydrolysis rather than synthesis under most conditions. Hence, the biosynthesis of peptide bonds requires an input of free energy. Nonetheless, peptide bonds are quite stable kinetically because the rate of hydrolysis is extremely slow; the lifetime of a reactivity in amino cyclization. acids. Some amino Homoserine can cyclize to acids are unsuitable for 35 proteins because 2.2 Primary of undesirable Structure peptide bond in aqueous solution in the absence of a catalyst approaches 1000 years. A series of amino acids joined by peptide bonds form a polypeptide chain, and each amino acid unit in a polypeptide is called a residue. A poly peptide chain has directionality because its ends are different : an a -amino group H R1 C O C OC – + +H N 3 +H N 3 H R2 C C O O C O– +H N 3 H R1 C O + H2O – C H N O R2 H Peptide bond FIGURE 2.13 Peptide-bond formation. The linking of two amino acids is accompanied by the loss of a molecule of water. 36 CH3 CHAPTER 2 Protein Composition and HC CH3 Structure O OH O H +H 3N C C H NC HH O HH C N H H2C H2C H H C N C C H2CH O C N H COC – O Tyr Gly Gly Phe Leu Amino terminal residue Carboxyl terminal residue theory of matter. Kilodalton (kDa) Kilodalton (kDa) A unit of mass equal to 1000 daltonsA unit of mass equal to 1000 daltons FIGURE 2.14 Amino acid sequences have direction. This illustration of the pentapeptide Tyr-Gly-Gly-Phe-Leu (YGGFL) shows the sequence from the amino terminus to the carboxyl terminus. This pentapeptide, Leu-enkephalin, is an opioid peptide that modulates the perception of pain. The reverse pentapeptide, Leu-Phe-Gly-Gly-Tyr (LFGGY), is a different molecule and has no such effects. is present at one end and an a -carboxyl group at the other. The amino end is taken to be the beginning of a polypeptide chain; by convention, the sequence of amino acids in a polypeptide chain is written starting with the amino-terminal residue. Thus, in the polypeptide Tyr-Gly-Gly-Phe Leu (YGGFL), tyrosine is the amino-terminal (Nterminal) residue and leucine is the carboxylterminal (C-terminal) residue (Figure 2.14). Leu PheGly-Gly-Tyr (LFGGY) is a different polypeptide, with Dalton different chemical properties. Dalton A unit of mass very nearly equal to A polypeptide chain consists of a regularly repeating that of a part, called the main chain or backbone, and a A unit of mass very nearly equal to that of a hydrogen atom. Named after John variable part, comprising the distinctive side chains (Figure 2.15). The polypeptide backbone is rich in Dalton hydrogen atom. Named after John Dalton hydrogen- bonding potential. Each residue contains (1766–1844), who developed the atomic a carbonyl group (C “ O), which is a good hydrogen(1766–1844), who developed the atomic theory of matter. bond acceptor, and, with the exception of proline, an consists of more than 27,000 amino acids. NH group, which is a good hydrogen-bond donor. Polypeptide chains made of small numbers of amino These groups interact with each other and with acids are called oligopeptides or simply peptides . functional groups from side chains to stabilize particular structures, as will be discussed in Section The mean molecular weight of an amino acid 2.3. residue is about 110 g mol 1 , and so the molecular Most natural polypeptide chains contain between 50 weights of most proteins are between 5500 and and 2000 amino acid residues and are commonly 220,000 g mol 1 . We can also refer to the mass of a referred to as proteins . The largest single protein, which is expressed in units of polypeptide known is the muscle protein titin , which CCO NH daltons; with a a mass of proteins, O C H HR1 N one dalton molecular 50,000 the linear HR3 O N C H is equal to weight of daltons, or polypeptide C N C one atomic 50,000 g 50 kDa chain is HR5 H 1 C mass unit. (kilodaltons C mol has H O C H C A protein ). In some N H O common cross-links are disul R2 cross-linked. The most R4 cysteine residues (Figure 2.16). The resulting unit of FIGURE 2.15 Components of a polypeptide chain. A polypeptide chain two linked cysteines is called cystine . Extracellular consists of a constant backbone (shown in black) and variable side chains (shown in green). proteins fide bonds, formed by the oxidation of a pair of often have several disulfide OH bonds, whereas intracel lular proteins usually lack cross-links derived from other C OH them. Rarely, nondisulfide side chains are present C N H in proteins. For example, collagen fibers in connective C N H2C C tissue are strengthened in this way, as are fibrin blood S H clots (Section 10.4). H2C H sequences S Cysteine + 2 e– +2 H Proteins have unique amino acid specified by genes (Figure 2.17). biochemistry because it showed C H S In 1953, Frederick CH2 Sanger determined H the amino acid This work is a sequence of insulin, a landmark in protein hormone for the first time that a protein has a precisely defined amino N Reduction + Oxidation H C CH2 C N S C OH by peptide bonds. This O H Cysteine accomplishment stimulated Cystine other scientists to carry out sequence studies of a wideknown. The striking variety of proteins. Currently, the complete amino FIGURE 2.16 Cross-links. The formation of a disulfide bond from two acid sequences of more than 2,000,000 proteins are cysteine residues is an oxidation reaction. acid sequence consisting only of L amino acids linked fact is that each protein has a unique, precisely defined amino acid sequence . The amino acid sequence of a protein is referred to as its primary structure . SS A chain Gly-Ile-Val-Glu-Gln-Cys-Cys-Ala-Ser-Val-Cys-Ser-Leu-Tyr-Gln-Leu-Glu-Asn-TyrCys-Asn 5 10 15 21 SS S S B chain Phe-Val-Asn-Gln-His-Leu-Cys-Gly-Ser-His-Leu-Val-Glu-Ala-Leu-Tyr-Leu-Val-Cys-Gly-Glu-Arg-Gly-Phe-Phe-Tyr-Thr-ProLys-Ala 5 10 15 20 25 30 FIGURE 2.17 Amino acid sequence of bovine insulin. A series of incisive studies in the late 1950s and early 1960s revealed that the amino acid sequences of proteins are determined by the nucleotide sequences of genes. The sequence of nucleotides in DNA specifies a com plementary sequence of nucleotides in RNA, which in turn specifies the amino acid sequence of a protein. In particular, each of the 20 amino acids of the repertoire is encoded by one or more specific sequences of three nucleotides (Section 4.6). Knowing amino acid sequences is important for several reasons. First, knowledge of the sequence of a protein is usually essential to elucidating its function (e.g., the catalytic mechanism of an enzyme). In fact, proteins with novel properties can be generated by varying the sequence of known proteins. Second, amino acid sequences determine the three-dimensional structures of proteins. The amino acid sequence is the link between the genetic message in DNA and the three-dimensional structure that performs a protein’s biological function. Analyses of relations between amino acid sequences and three-dimensional structures of proteins are uncovering the rules that govern the folding of polypeptide chains. Third, alterations in amino acid sequence can lead to abnormal protein function and disease. Severe and sometimes fatal diseases, such as sickle-cell anemia (Chapter 7) and cystic fibrosis, can result from a change in a single amino acid within a protein. Fourth, the sequence of a protein reveals much about its evolution ary history (Chapter 6). Proteins resemble one another in amino acid sequence only if they have a common ancestor. Consequently, molecular events in evolution can be traced from amino acid sequences; molecular paleontology is a flourishing area of research. atom and CO group of the first amino acid and the NH group and a -carbon atom of the second amino acid. The nature of 38 CHAPTER 2 Protein Composition and Structure H Polypeptide chains are flexible yet conformationally restricted Examination of the geometry of the protein the chemical bonding within a peptide accounts for backbone reveals several important features. First, the bond’s planarity. The bond resonates between a the peptide bond is essentially planar (Figure 2.18). single bond and a double bond. Because of this Thus, for a pair of amino acids linked by a peptide partial double-bond character, rotation about this bond, six atoms lie in the same plane: the a -carbon bond is prevented and thus the conforma backbone is constrained. Cα N C αC H tion of the peptide C C O FIGURE 2.18 Peptide bonds are planar. In a pair of linked amino acids, six O N C atoms (C , C, O, N, H, and C ) lie in a plane. Side chains are shown as green balls. Peptide-bond resonance structures H C C N+ C O– The partial double-bond character is also expressed in the length of the bond between the CO and the NH groups. As shown in Figure 2.19, the C}N distance in a peptide bond is typically 1.32 Å, which is between the values expected for a C}N single 37 bond (1.49 Å) and a C “ N Cα to form tightly trans over cis common cis packed globularcan be peptide bonds structures. explained by the are X}Pro fact that steric linkages. Such Two bonds show configurations clashes are possible for between groups less preference for the trans a planar peptide configuration bond. In the because the trans nitrogen of configuration, H proline is the two a bonded to two carbon atoms tetrahedral are on opposite N carbon atoms, sides of the 1.0 Å peptide bond. attached to the limiting the Cα 1.45 Å In the cis a -carbon atoms steric 1.51 Å differences configuration, hinder the double bond (1.27 Å). Finally, these groups C formation of the between the trans and cis are on the same cis the peptide 1.32 Å forms (Figure side of the pep configuration bond is tide bond. uncharged, but do not arise 2.21). Almost all allowing In contrast with in the trans peptide bonds inconfiguration polymers of the peptide proteins are amino acids (Figure 2.20). bond, the bonds trans . This linked by By far the most between the peptide bonds preference for 1.24 Å O two adjacent rigid peptide units can rotate about these bonds, taking on various amino group and the a -carbon atom and orientations. This freedom of rotation about between the a -carbon atom and the two carbonyl group are pure single bonds. The FIGURE 2.19 Typical bond lengths within a peptide unit. The peptide unit is shown in the trans configuration. bonds of each amino acid allows proteins to fold in many different ways . The rotations about these bonds can be specified by Trans Cis FIGURE 2.20 Trans and cis peptide bonds. The trans form is strongly favored because of steric clashes, indicated by the orange semicircles, that arise in the cis form. 39 2.2 Primary Structure Trans Cis FIGURE 2.21 Trans and cis X–Pro bonds. The energies of these forms are similar to one another because steric clashes, indicated by the orange semicircles, arise in both forms. (B) (C) RH RH OH (A) C C C C N C N C N HH O HR O = +85° = −80° fold into well defined structures is remarkable thermodynamically. An unfolded polymer exists as a acid in a polypeptide can be adjusted by rotation about two single bonds. (A) Phi ( ) is the angle of rotation about the bond between the nitrogen and the a- random coil: each copy of an unfolded polymer will carbon atoms, whereas psi ( ) is the angle of rotation about the bond between thehave a differ ent conformation, yielding a mixture of a-carbon and the carbonyl carbon atoms. (B) A view down the bond between many possible conformations. The favorable the nitrogen and the a-carbon atoms, showing how is measured. (C) A view entropy associated with a mixture of many down the bond between the a-carbon and the carbonyl carbon atoms, showing conformations opposes folding and must be how is measured. overcome by interactions favoring the folded form. Thus, highly flexible polymers with a large number of possible con formations do not fold into unique torsion angles (Figure 2.22). The angle of rotation structures. The rigidity of the peptide unit and the about the bond between the nitrogen and the a carbon atoms is called phi ( ). The angle of rotation restricted set of allowed f and c angles limits the number of structures accessible to the unfolded form about the bond between the a -carbon and the sufficiently to allow protein folding to take place. carbonyl carbon atoms is called psi ( ) . A clockwise rotation about either bond as viewed from the nitrogen atom toward the a -carbon atom or from the a -carbon atom toward the carbonyl group corresponds to a positive value. The and angles determine the path of the polypeptide chain. Are all combinations of and possible? Gopalasamudram Ramachandran recognized that many combinations are forbidden because of steric collisions between atoms. The allowed values can be visualized on a two-dimensional plot called a Ramachandran plot (Figure 2.23). Three quarters of the possible ( , ) combinations are excluded simply by local steric clashes. Steric exclusion, the fact that two atoms cannot be in the same place at the same time, can be a powerful organizing principle. Torsion angle The ability of biological polymers such as proteins to A measure of the rotation about a bond, usually taken to lie between 2180 and 1180 FIGURE 2.22 Rotation about bonds in a polypeptide. The structure of each amino degrees. Torsion angles are sometimes called dihedral angles. +180 40 CHAPTER 2 Protein Composition and Structure 120 60 0 −60 −120 ( = 90°, = −90°) −180 −60 −120 −180 600 120 +180 Disfavore d 2.3 Secondary Structure: Polypeptide Chains Can Fold into Regular Structures Such As the Alpha Helix, the Beta Sheet, and Turns and Loops Can a polypeptide chain fold into a regularly repeating structure? In 1951, Linus Pauling and Robert Corey proposed two periodic structures called the helix (alpha helix) and the pleated sheet (beta pleated sheet). Subsequently, other structures such as the turn and omega ( V ) loop were identified. Although not periodic, these common turn or loop structures are well defined and contribute with a helices and b sheets to form the final protein structure. Alpha helices, b strands, and turns are formed by a regu lar pattern of hydrogen bonds between the peptide N}H and C “ O groups of amino acids that are near one another in the linear sequence . Such folded segments are called secondary structure . The alpha helix is a coiled structure stabilized by intrachain hydrogen bonds In evaluating potential structures, Pauling and Corey considered which con formations of peptides were sterically allowed and which most fully exploited the hydrogen-bonding capacity of the backbone NH and CO groups. The first of their proposed structures, the helix, is a rodlike structure (Figure 2.24). A tightly coiled backbone forms the inner part of the rod and the side chains extend outward in a helical array. The a helix is stabilized by hydrogen bonds between the NH and CO groups of the main chain. In par ticular, the CO group of each amino acid forms a hydrogen bond with the NH group of the amino acid that is situated four residues ahead in the sequence (Figure 2.25). Thus, except for amino acids near the ends of an a helix, all the main-chain CO and NH groups are hydrogen bonded. Each resi due is Screw sense related to the next one by a rise, also called Screw sense Describes the direction in which a helical translation, of 1.5 Å along the helix axis and a Describes the direction in which a helical structure rotates with respect to its axis. If, rotation of 100 degrees, which gives 3.6 amino acid structure rotates with respect to its axis. resi dues per turn of helix. Thus, amino acids spaced If, viewed down the axis of a helix, three and four apart in the sequence are spatially the chain turns quite close to one another in an a helix. In contrast, viewed down the axis of a helix, the chain turns in a clockwise direction, it has a amino acids spaced two apart in the sequence are right-handed in a clockwise direction, it has a right- situated on opposite sides of the helix and so are handed screw sense. If the turning unlikely to make contact. The pitch of the a helix is is counterclockwise, the length of one complete turn along the helix axis screw sense. If the turning is counterclockwise, the screw sense is left-handed. and is equal to the product of the rise (1.5 Å) and the the screw sense is left-handed. FIGURE 2.23 A Ramachandran plot showing the number of residues per turn (3.6), or 5.4 Å. The values of and . Not all and values are possible without collisions between screw atoms. The most favorable regions are shown in dark green; borderline regions are shown in light green. The structure on the right is disfavored because of steric clashes. (A) (B) (C) (D) handed +180 120 60 0 41 2.3 Secondary Structure Left- −60 FIGURE 2.24 Structure of the a helix. (A) A ribbon depiction shows the a- carbon atoms and side chains (green). (B) A side view of a ball-and-stick version depicts the hydrogen −120 helix (very rare) bonds (dashed lines) between NH and CO groups. (C) An end view shows the coiled backbone as the inside of the helix and the side chains (green) projecting outward. (D) A helix (common) −120 −180 Right-handed space-filling view of part interior core of the helix. −60 C shows the tightly packed 600 120 −180 +180 FIGURE 2.26 Ramachandran plot for O HO Ri+2 H i+4 H H H R Ri N C H O Ri+1 helices. Both right- and left-handed helices lie in regions of allowed conformations in the HO C N C C N C C N HO OH HR Ri+5 H i+3H C C N C C N (A) (B) FIGURE 2.25 Hydrogen-bonding scheme for an a helix. In the a helix, the CO C C However, all a helicesare rightRamachandr essentially in proteins handed. an plot. are in a helices (Figure 2.28). Indeed, about 25% of all soluble proteins are composed of a helices sense of an a helix can be right-handed (clockwise) connected by loops and turns of the polypeptide or left-handed (counter clockwise). The chain. Single a helices are usually less than 45 Å Ramachandran plot reveals that both the rightlong. Many proteins that span biological membranes handed and the left-handed helices are among also contain a helices. allowed conformations (Figure 2.26). However, righthanded helices are energetically more favorable because there is less steric clash between the side chains and the backbone. Essentially all helices found in proteins are right-handed. In schematic representations of proteins, a helices are depicted as twisted ribbons or rods (Figure 2.27). Not all amino acids can be readily accommodated in an a helix. Branching at the b -carbon atom, as in valine, threonine, and isoleucine, tends to destabilize a helices because of steric clashes. Serine, aspartate, and asparagine FIGURE 2.27 Schematic views of a helices. (A) A ribbon depiction. (B) A also tend to disrupt a helices because their side cylindrical depiction. chains contain hydrogen-bond donors or acceptors in close proximity to the main chain, where they compete for main-chain NH and CO groups. Proline also is a helix breaker because it lacks an NH group and because its ring structure prevents it from assuming the value to fit into an a helix. The a helical content of proteins ranges widely, from none to almost 100%. For example, about 75% of the residues in ferritin, a protein that helps store iron, FIGURE 2.28 A largely a-helical protein. Ferritin, an iron-storage protein, is group of residue i forms a hydrogen bond with the NH group of residue i 1 4. built from a bundle of a helices. [Drawn from 1AEW.pdb.] is composed of two or Pauling and Corey proposed another periodic more polypeptide chains called strands. A b structural motif, which −60 they named the pleated strand is almost fully sheet ( b because it was extended rather than −120 the second structure that being tightly coiled as in the a helix. A range of they elucidated, the a extended structures are helix having been the −120 −180 first). The b pleated sheet sterically allowed (Figure 2.29). Beta strands (or, more simply, the b Beta sheets are stabilized by The distance between sheet) differs markedly hydrogen bonding between adjacent amino acids from the rodlike a helix. It polypeptide strands along a b strand is approxi 600 distance of 1.5 Å opposite directions +180 120 along an a helix. chains of adjacent (Figure 2.30). mately 3.5 Å, in The side amino acids point in contrast with a 0 +180 120 60 −180 −60 FIGURE 2.29 Ramachandran plot for b strands. The red area shows the sterically allowed conformations of extended, b-strand-like structures. 7Å FIGURE 2.30 Structure of a b strand. The side chains (green) are alternately above and below the plane of the strand. A b sheet is formed by linking two or more b strands lying next to one another through hydrogen bonds. Adjacent strands in a b sheet can run in opposite directions (antiparallel b sheet) or in the same direction (parallel b sheet). In the antiparallel arrangement, the NH group and the CO group of each amino acid are respectively hydrogen bonded to the CO group and the NH group of a partner on the adjacent chain (Figure 2.31). In the parallel arrangement, the hydrogen-bonding scheme is slightly more complicated. For each amino acid, the NH group is hydrogen bonded to the CO group of one amino acid on the adjacent strand, whereas the CO group is hydrogen bonded to the NH group on the amino acid two residues farther along the chain (Figure 2.32). Many strands, typically 4 or 5 but as many as 10 or more, can come together in b sheets. Such b sheets can be purely antiparal lel, purely parallel, or mixed (Figure 2.33). FIGURE 2.31 An antiparallel b sheet. Adjacent b strands run in opposite directions, as indicated by the arrows. Hydrogen bonds between NH and CO groups connect each amino acid to a single amino acid on an adjacent strand, stabilizing the structure. 42 2.3 Secondary Structure FIGURE 2.32 A parallel b sheet. Adjacent b strands run in the same direction, as indicated by the arrows. Hydrogen bonds connect each amino acid on one strand with two different amino acids on the adjacent strand. FIGURE 2.33 Structure of a mixed b sheet. The arrows indicate directionality of each strand. In schematic representations, b strands are usually depicted by broad arrows pointing in the direction of the carboxyl-terminal end to indicate the type of b sheet formed—parallel or antiparallel. More structurally diverse than a helices, b sheets can be almost flat but most adopt a somewhat twisted shape (Figure 2.34). The b sheet is an important structural element in many proteins. For example, fatty acid-binding proteins, important for lipid metabolism, are built almost entirely from b sheets (Figure 2.35). (A) (B) FIGURE 2.35 A protein rich in b FIGURE 2.34 A schematic twisted b sheet. (A) A schematic model. (B) The schematic view rotated by 90 degrees to illustrate the twist more clearly. sheets. The structure of a fatty acid binding protein. [Drawn from 1FTP.pdb.] 44 CHAPTER 2 Protein Composition and Structure i + 1i + 2 i+3 i FIGURE 2.36 Structure of a reverse turn. The CO group of residue i of the polypeptide chain is hydrogen bonded to the NH group of residue i 1 3 to stabilize the turn. FIGURE 2.37 Loops on a protein surface. A part of an antibody molecule has surface loops (shown in red) that mediate interactions with other molecules. [Drawn from 7FAB.pdb.] Polypeptide chains can change direction by making reverse Most proteins have compact, globular shapes owing to reversals in the direction of their polypeptide chains. Many of these reversals are accom plished by a common structural element called the reverse turn (also known as the turn or hairpin turn ) , illustrated in Figure 2.36. In many reverse turns, the CO group of residue i of a polypeptide is hydrogen bonded to the NH group of residue i 1 3. This interaction stabilizes abrupt changes in direction of the polypeptide chain. In other cases, more-elaborate structures are responsible for chain reversals. These structures are called loops or sometimes loops (omega loops) to suggest their overall shape. Unlike a helices and b strands, loops do not have regular, periodic structures. Nonetheless, loop structures are often rigid and well defined (Figure 2.37). Turns and loops invariably lie on the surfaces of proteins and thus often participate in interactions between proteins and other molecules. Fibrous proteins provide structural support for cells and tissues Special types of helices are present in the two proteins a-keratin and collagen. These proteins form long fibers that serve a structural role. a Keratin, which is an essential component of wool, hair, and skin, con sists of two right-handed a helices intertwined to form a type of left-handed superhelix called an -helical coiled coil . a -Keratin is a member of a superfam ily of proteins referred to as coiled-coil proteins (Figure 2.38). In these proteins, two or more a helices can entwine to form a very stable structure, which can have a length of 1000 Å (100 nm, or 0.1 m m) or more. There are approximately 60 members of this family in humans, including intermediate filaments, pro teins that contribute to the cell cytoskeleton (internal scaffolding in a cell), and the muscle proteins myosin and tropomyosin (Section 35.2). Members of this family are characterized by a central region of 300 amino acids that contains imperfect repeats of a sequence of seven amino acids called a heptad repeat . The two helices in a -keratin associate with each other by weak interactions such as van der Waals forces and ionic interactions. The left-handed supercoil alters the two right-handed a helices such that there are 3.5 residues per turn instead of 3.6. Thus, the pattern of side-chain interactions can be repeated every seven residues, forming the heptad repeats. Two helices with such repeats are able to interact with one another if the repeats are complementary (Figure 2.39). For example, the repeating residues may be hydrophobic, allow ing van der Waals interactions, or have opposite charge, allowing ionic interac tions. In addition, the two helices may be linked by disulfide bonds formed by neighboring cysteine residues. The bonding of the helices accounts for the physical properties of wool, an example of an a -keratin. Wool is extensible and can be stretched to nearly twice its length because the a helices stretch, breaking (A) FIGURE 2.38 An a-helical coiled coil. (A) Space-filling model. (B) Ribbon diagram. The two helices wind around one another to form a superhelix. Such structures are found in many proteins, including keratin in hair, quills, claws, and horns. [Drawn from 1C1G.pdb.] (B) Leu the weak interactions strand are absent. between neighboring Instead, the helix is helices. However, the stabilized by steric covalent disulfide bonds repulsion of the pyrrolidine resist breakage and return rings of the proline and Leu the fiber to its original hydroxyproline resi dues state once the stretching (Figure 2.41). The force is released. The pyrrolidine rings keep out number of disulfide bond of each other’s way when Leu cross-links further defines the polypeptide chain the fiber’s properties. Hair assumes its helical form, and wool, having fewer which has about three resi cross-links, are flex dues per turn. Three strands wind around one Leu ible. Horns, claws, and another to form a hooves, having more superhelical cable that is cross-links, are much harder. A different type stabilized by hydrogen bonds between strands. of helix is present in The hydrogen collagen, the most abundant protein of mammals. Collagen is Leucine (Leu) residue the main fibrous component of skin, bone, tendon, cartilage, and teeth. This extracellular protein is a rod-shaped molecule, about 3000 Å long and only 15 Å in diameter. It contains three Leu helical polypeptide chains, each nearly 1000 residues long. Glycine appears at every third residue in the amino acid Leu sequence, and the sequence glycine-prolinehydroxyproline recurs frequently (Figure 2.40). Leu Hydroxyproline is a derivative of proline that C C N N has a hydroxyl group in place of one of the hydrogen atoms on the pyrrolidine ring. The collagen helix has properties different from those of the a helix. Hydrogen bonds within a Gly bonds form between the peptide NH groups of glycine residues and the CO groups of residues on the other chains. The hydroxyl groups of hydroxypro line residues also participate in hydrogen bonding. ProPro 13 FIGURE 2.39 Heptad repeats in a coiled-coil protein. Every seventh residue in glycine be present at every third position on each each helix is leucine. The two helices are held together by van der Waals strand (Figure 2.42A). The only residue that can fit in interactions an interior position is glycine . The amino acid primarily between the leucine residues. [Drawn from 2ZTA.pdb.] residue on either side of glycine is located on the outside of the cable, where there is room for the bulky rings of proline and hydroxyproline residues (Figure 2.42B). Pro -Gly-Pro-Met-Gly-Pro-Ser-Gly-Pro-Arg 22 -Gly-Leu-Hyp-Gly-Pro-Hyp-Gly-Ala-Hyp 31 -Gly-Pro-Gln-Gly-Phe-Gln-Gly-Pro-Hyp 40 -Gly-Glu-Hyp-Gly-Glu-Hyp-Gly-Ala-Ser 49 -Gly-Pro-Met-Gly-Pro-Arg-Gly-Pro-Hyp 58 -Gly-Pro-Hyp-Gly-Lys-Asn-Gly-Asp-Asp ProGly FIGURE 2.41 Conformation of a single strand of a collagen triple helix. FIGURE 2.40 Amino acid sequence of a part of a collagen chain. Every The inside of the triple-stranded helical cable is very third residue is a glycine. Proline and hydroxyproline (Hyp) also are abundant. crowded and accounts for the requirement that (A) (B) G G FIGURE 2.42 Structure of the protein collagen. (A) Space filling model of collagen. Each strand is shown in a different G color. (B) Cross section of a model of collagen. Each strand is hydrogen bonded to the other two strands. The a-carbon atom of a glycine residue is identified by the letter G. Every third residue must be glycine because there is no space in the center of the helix. Notice that the pyrrolidine rings of the proline residues are on the outside. Myoglobin is an extremely compact molecule . Its CHAPTER 2 Protein Composition and Structure overall dimensions are 45 3 35 3 25 Å, an order of The importance of the positioning of glycine inside magnitude less than if it were fully stretched out the triple helix is illustrated in the disorder (Figure 2.43). About 70% of the main chain is folded osteogenesis imperfecta, also known as brittle bone into eight a heli disease. In this condition, which can vary from mild toces, and much of the rest of the chain forms turns very and loops between helices. The folding of the main severe, other amino acids replace the internal glycinechain of myoglobin, like that of most other pro teins, residue. This replace ment leads to a delayed and is complex and devoid of symmetry. The overall improper folding of collagen. The most serious course of the poly peptide chain of a protein is symptom is severe bone fragility. Defective collagen referred to as its tertiary structure . A unifying in the eyes causes the whites of the eyes to have a principle emerges from the distribution of side chains. blue tint (blue sclera). Strikingly, the inte rior consists almost entirely of nonpolar residues such as leucine, valine, methionine, and phenylalanine (Figure 2.44). Charged residues such as aspartate, glutamate, 2.4 Tertiary Structure: Water-Soluble Proteins Fold into Compact Structures with Nonpolar Cores lysine, and arginine are absent from the inside of myo globin. The only polar residues inside are two Let us now examine how amino acids are grouped histidine residues, which play critical roles in binding iron and oxygen. The outside of myoglobin, on the together in a complete protein. X-ray crystallographic and nuclear magnetic resonance (NMR) studies (Section 3.5) have revealed the detailed three-dimensional structures of thousands of proteins. We begin here with an examination of myoglobin, the first protein to be seen in atomic detail. Myoglobin, the oxygen storage protein in muscle, is a single polypeptide chain of 153 amino acids (Chapter 7). The capacity of myoglobin to bind oxygen depends on the presence of heme, a nonpolypeptide prosthetic (helper) group consisting of protoporphyrin IX and a central iron atom. 46 45 (B) Heme group (A) Heme group Iron atom FIGURE 2.43 Three-dimensional structure of myoglobin. (A) A ribbon diagram shows that the protein consists largely of a helices. (B) A space-filling model in the same orientation shows how tightly packed the folded protein is. Notice that the heme group is nestled into a crevice in the compact protein with only an edge exposed. One helix is blue to allow comparison of the two structural depictions. [Drawn from 1A6N.pdb.] model of myoglobin with hydrophobic amino acids shown in yellow, charged amino acids shown in blue, and others shown in white. Notice that the surface of the molecule has many charged amino acids, as well as some hydrophobic amino acids. (B) In this cross-sectional view, notice that mostly hydrophobic amino acids are found on the inside of the structure, whereas the charged amino acids are found on the protein surface. [Drawn from 1MBD.pdb.] other hand, consists of both polar and nonpolar residues. The space-filling model shows that there is very little empty space inside. This contrasting distribution of polar and nonpolar residues reveals a key facet of protein architecture. In an aqueous environment, protein folding is driven by the strong tendency of hydrophobic residues to be excluded from water. Recall that a system is more thermodynamically stable when hydrophobic groups are clustered rather than extended into the aqueous surroundings (p. 9). The polypeptide chain therefore folds so that its hydrophobic side chains are buried and its polar, charged chains are on the surface . Many a helices and b strands are amphipathic; that is, the a helix or b strand has a hydrophobic face, which points into the protein interior, and a more polar face, which points into solution. The fate of the main chain accompanying the “the exceptions that prove the rule” because they hydrophobic side chains is important, too. An have the reverse distribution of hydrophobic and unpaired peptide NH or CO group markedly prefers hydrophilic amino acids. For example, consider water to a nonpolar milieu. The secret of burying a porins, proteins found in the outer membranes of segment of main chain in a hydrophobic many bacteria (Figure 2.45). Membranes are built environment is to pair all the NH and CO groups by largely of hydropho hydrogen bonding. This pairing is neatly bic alkane chains (Section 12.2). Thus, porins are accomplished in an a helix or b sheet. Van der cov ered on the outside largely with hydrophobic Waals interactions between tightly packed hydrocar residues that interact with the neighboring alkane bon side chains also contribute to the stability of pro chains. In contrast, the center of the protein contains many charged and polar amino acids that teins. We can now understand why the set of 20 surround a water-filled chan nel running through the amino acids contains several that differ subtly in size and shape. They provide a palette from which middle of the protein. Thus, FIGURE 2.45 “Inside out” amino acid distribution in porin. The outside of to choose to fill the interior of a protein neatly and porin (which contacts hydrophobic groups thereby maxi mize van der Waals interactions, which in membranes) is covered largely with hydrophobic residues, whereas the require inti mate contact. center includes a water-filled channel lined with charged and polar amino Some proteins that span biological membranes are acids. [Drawn from 1PRN.pdb.] hydrophobic environments, they are “inside out” relative to proteins that function in aqueous solution. Water-filled hydrophilic channel Largely hydrophobic exterior because porins function in 47